AI now writes almost half of all code. Vetting a dev agency just got harder.

Key takeaways
Everyone is fast now
Google's DORA research found AI-assisted teams ship about 20% more pull requests per developer, with 23.5% more incidents per change. Speed stopped being a way to compare agencies. Discipline is where the spread lives.
The data on unreviewed AI code is consistent
GitClear measured code churn doubling since 2022, duplicated blocks up eightfold, refactoring down 70%. Veracode found a security flaw in 45% of AI coding tasks. Fast and clean are different claims.
Ask for maintenance stories, not portfolios
A portfolio shows launches. A project the agency has maintained past its first year shows code that survived. That is the strongest single signal you can ask for.
Cheap quotes defer cost, they don't remove it
AI compressed the cost of writing code, not the cost of reviewing and testing it. If a low quote cut those to hit the price, you pay for them later, at rescue rates.
Two numbers frame the problem. GitHub's telemetry shows Copilot now writes about 46% of the code produced by the developers who use it, up from 27% when it launched commercially in 2022. Gartner expects 60% of all new code to be AI-generated by the end of this year. Whichever agency you are about to hire, AI is writing a large share of your product. That part is settled.
What is not settled is how you pick the agency. The signals founders have always used to compare shops, portfolio polish, demo speed, day rate, were never great. Now they are close to useless, because AI compressed exactly the parts of the work those signals used to measure.
Why the old signals stopped working
A polished demo used to mean something. Getting a clean UI, working auth, and a believable flow in front of a client took weeks of competent work, so the demo doubled as evidence of a competent team. Today the same demo is evidence of a twenty-dollar-a-month subscription. Any two-person shop can scaffold it in a day. Plenty do.
Velocity flattened out too. Google's DORA research found that teams adopting AI ship roughly 20% more pull requests per developer, and that incidents per change rose 23.5% alongside. More code goes out the door, and each piece of it is slightly more likely to break something. When every agency you talk to is fast, speed is no longer a thing you can select on.
Price stopped discriminating from the other direction. Quotes for the same scope now range from a few thousand dollars to twenty times that, and the cheap end is not necessarily lying. They really can build it at that price. What the quote doesn't tell you is the state of the code afterward, and who pays for that state later. Usually you.
What unreviewed AI code looks like six months in
The research here is blunt. GitClear, which analyzed over 600 million code changes, found that code churn, meaning lines rewritten or thrown away within weeks of being written, has roughly doubled since 2022. Duplicated code blocks rose eightfold in a single year. Refactoring, the unglamorous consolidation work that keeps a codebase maintainable, fell about 70%. Code is being produced faster and organized less.
Security tells the same story. Veracode ran more than 100 models through 80 coding tasks and got a security flaw in 45% of the results. When a task could be solved in a secure or an insecure way, the models regularly picked the insecure one. For cross-site scripting specifically, the failure rate hit 86%.
Here is the part I find telling: the people closest to these tools trust them least. In Stack Overflow's 2025 survey, 46% of developers said they don't trust the accuracy of AI output, against 33% who do, and the most experienced developers were the most skeptical of all. Almost half called debugging AI-generated code a time sink. Adoption kept climbing anyway, to 84%. Developers didn't stop using AI. They stopped assuming it's right.
None of this makes AI-assisted development a mistake. We use these tools daily and the speed gain is real. But the gain arrives now and the debt arrives in month three, which is why rescuing AI-built apps became its own service line with its own price list. The agencies generating that rescue work and the agencies doing the rescuing look identical in a pitch meeting. Both are fast. The difference sits in process you can't see from a demo.
Seven questions that still separate agencies
You can't audit a codebase before you've hired the people who'd write it. You can listen closely to how an agency talks about process. On these questions, disciplined and undisciplined shops give visibly different answers.
- How do you use AI in development? Both extreme answers are bad. "We don't" is either untrue or a sign you're paying artisan prices for commodity work. "AI does most of it" with nothing after is worse. The answer you want names actual tools and then, unprompted, describes how the output gets reviewed.
- Who reviews AI-generated code, and how? Listen for mechanics: pull requests, a named reviewer, review before merge rather than after launch. "The AI checks itself" or "our seniors keep an eye on things" both translate to: nobody reviews it.
- What does testing look like on a real project? Ask to see the CI pipeline from something they shipped. An agency moving at AI speed either has serious automated testing or is shipping that 45% security statistic straight to your production.
- Can you show me a project you've maintained for over a year? Portfolios show launches. The GitClear numbers describe what happens after launch: churn, duplication, code nobody consolidates. A shop still holding maintenance contracts from 2024 has code that survived contact with reality. Strongest signal on this list.
- Where did the review and testing time go in this quote? If the price came in at half what you expected, something got cut. Sometimes AI really did cut the hours and the quote is honest. Ask them to walk the line items and watch whether review and testing survived the discount.
- Do you run security scanning, and where? The answer should include static analysis inside the CI pipeline, not a one-off audit at the end. With nearly half of raw AI output carrying a flaw, scanning as you go is the difference between catching it on Tuesday and hearing about it from a customer.
- What does handover look like if we part ways? Duplicated, churning code is hostile to any team that didn't write it. You want documentation, deployment access, and a codebase another agency could pick up. If they get vague here, the mess may be the retention strategy.
Notice that none of these require you to read code. They require the agency to describe its own process in specifics, which disciplined shops do easily because they're describing what they did this morning. Undisciplined shops will steer the conversation back to their tools. Tools are the one thing that no longer separates anyone.
Frequently asked questions
Should I avoid agencies that use AI to write code?
No, and you couldn't if you tried. GitHub's own telemetry puts AI at 46% of the code written by developers using Copilot, and Gartner expects 60% of all new code to be AI-generated by the end of 2026. The useful filter is not whether an agency uses AI but whether they can describe, in specifics, how its output gets reviewed and tested before it ships.
Is AI-generated code really less secure?
Unreviewed, yes. Veracode tested more than 100 models on 80 coding tasks and found a security flaw in 45% of the results; when a task could be solved securely or insecurely, models regularly picked the insecure route. With static analysis in the pipeline and human review before merge, those flaws get caught. That is the setup you should ask for.
What is the single best question to ask a dev agency in 2026?
Ask to see a project they have maintained for more than a year, and what changed in it recently. Portfolios show launch day. Maintenance contracts show whether the code held up after it. An agency whose clients keep paying them to evolve old codebases is holding evidence a demo can't fake.
Why are agency quotes so far apart for the same scope?
Because AI collapsed the cost of producing code while the cost of reviewing, testing, and securing it stayed put. Quotes at the low end usually cut the second category. The build still ships. The bill for what was skipped arrives in a few months, and rescue work on AI-built apps now runs from a few thousand dollars for an audit to six figures for a rebuild.
Related posts
ChatGPT's app store is six months old. Should you build there yet?
900 million people use ChatGPT every week, and it now has an app SDK, a directory, and its first 300 apps. The distribution pitch is real. The traffic, so far, mostly isn't.
What it costs to rescue a vibe-coded app in 2026
Fixing AI-built apps is now its own service line, and the quotes run from a $3,000 audit to a half-million rebuild. Here are the real tiers, and the five signals that tell you which one you're in.