BrightStar IT Management & Consulting
AI-strategie10 min leestijd

AI writes the code. Who still knows the system?

R
Raôul Zon
Technology leader, interim & advisory · 11 October 2026
Dark automated factory floor with robotic arms and a single lit control deskIn het Engels

AI coding agents are now standard issue in most engineering teams. The question boards ask has moved on from "should we use them?" to "where is the productivity we were promised?"

The honest answer: individual developers do produce more code, but most organisations are not shipping more value. The gain gets lost somewhere between the keyboard and the customer.

That gap is not a tooling problem. It is a leadership problem, and it changes how you should run, staff and budget your engineering organisation. This article looks at what the evidence actually says, where the bottleneck moves to, and what technology leaders should do about it.

Feeling faster is not the same as being faster

Ask developers and they will tell you AI makes them more productive. More than 80% said so in Google's 2025 DORA research.

Measure it and the picture is less clear. When METR tested experienced developers on real work in 2025, they were 19% slower with AI, while believing they were 20% faster. The tools have improved since, and METR's 2026 follow-up points to a modest speedup. Real, but nothing like the doubling on the vendor slides.

The lesson for leaders: self-reported productivity is not evidence. If your business case rests on a developer survey, it rests on perception. Developers sense this themselves; in the 2025 Stack Overflow survey, more of them distrusted AI output than trusted it.

The bottleneck moves downstream

Writing code was never the main constraint in a mature engineering organisation. Understanding the problem, reviewing, testing, integrating and releasing safely were. AI speeds up the part that was already fastest and pushes more volume into everything behind it.

Data from more than 10,000 developers, analysed by Faros AI, shows the pattern. Heavy AI users merged almost twice as many pull requests, but review time nearly doubled too. At company level, delivery did not improve.

DORA puts it well: AI is an amplifier. It magnifies the strengths of high-performing organisations and the dysfunctions of struggling ones.

In practice, this shows up in three places:

  • Your senior engineers become the bottleneck. They review larger changes, written faster, by people (and agents) who understand them less.
  • Your test and release pipeline gets stress-tested. Weak automated testing was survivable at human speed. At AI speed it becomes your incident rate.
  • Your architecture shows its seams. Agents work well inside clean boundaries and well-documented API's. In a tangled monolith, they produce plausible code that breaks something three modules away.

This is why AI coding tools reward the organisations that need them least. If your delivery is stable, your boundaries clear and your pipeline automated, AI makes you faster. If not, it makes you busier.

When agents review agents, who still knows the system?

The obvious answer to the review bottleneck is to let agents review as well. Many teams already do: one agent writes, another reviews, a third tests, and an orchestrator ties it all together. Humans write the spec at the start and look at the result at the end.

That opens a new gap. The spec says what the system should do. The reviewed output says it passed the checks. What sits in between, the detailed knowledge of how the system actually works, why it was built that way and where the edge cases hide, no longer passes through anyone's head.

Taken to its end point, this is the "dark factory", named after factories that run with the lights off because no people work there. Security company StrongDM calls its version a software factory, and its charter, published in February 2026, has two rules: code must not be written by humans, and code must not be reviewed by humans. In their own words, validation replaces code review. A Stanford Law commentary on it asks the obvious question: when that software fails, who is accountable, and who is still able to read the code to find out why?

This is not an argument against it. For well-bounded components with strong automated tests, a dark factory can be a rational choice. But it should be a deliberate decision at leadership level, not a place you drift into one convenience at a time. Three questions help:

  • Which systems can we run without anyone who understands the code? A campaign microsite is not a payments engine.
  • What replaces human review as the safety net? Specs and tests become the real source code, so they deserve the rigour code used to get.
  • Who is accountable when it goes wrong? And can they explain it to a customer, an auditor or a regulator?

The junior pipeline is the hidden cost

The tempting conclusion is simple: agents do junior work, so hire fewer juniors. Many companies already act on it.

Stanford research shows it is already happening. In the US, employment of developers aged 22 to 25 fell by around 13% compared with less AI-exposed jobs, while employment of experienced developers held steady. The explanation is telling: AI replaces what juniors learned at school, not what seniors learned on the job.

The problem is that craftmanship has only one source: years of doing the work. Juniors became seniors by writing boilerplate, fixing bugs, getting their pull requests torn apart in review and breaking things in production environments. That is exactly the work agents now take over.

Cut the junior intake today and you save some salary cost. In five years you have no mid-level engineers, in eight no seniors, and the people who could judge whether an agent's output is safe to ship are retiring or become very expensive to hire.

The smarter move is to keep hiring juniors and redesign how they learn. Pair them with seniors on review, not just on writing code. Make them explain and defend agent-generated changes. Treat the ability to judge code as the skill you are training, because that is the skill AI makes scarce.

Team shape and sourcing need a rethink

If producing code gets cheaper, what you pay for changes. The scarce resource becomes judgement: knowing what to build, how it fits the architecture and whether it is safe to release. That has consequences for how you shape and source teams.

Team shape. Expect smaller teams with a higher share of senior engineers, plus a stronger platform function that owns the pipeline, test automation and guardrails the agents run inside. The classic pyramid with many juniors at the base becomes a diamond. That is fine, as long as you keep feeding the base (see above).

Nearshoring and offshoring. Most sourcing models were built on one idea: buy coding capacity at a lower rate. When coding itself gets cheaper, the rate difference matters less and the coordination cost matters more. That plays out differently depending on which side of the model you are on.

If you buy capacity from a partner, a supplier who delivers volume but leaves review, integration and accountability with your in-house seniors now adds load to your bottleneck. The partners worth keeping deliver outcomes, own quality end to end and bring their own senior judgement. Contracts priced on headcount deserve a fresh look.

If you run your own delivery centres, as most large agencies and SI's do, the pressure is on your business model. Those centres are typically built as a pyramid: many juniors producing code, a thin senior layer onshore close to the client. AI compresses exactly the work at the base of that pyramid, and with it the billable hours. Selling FTE's at a blended rate gets harder to defend when clients know an agent writes much of the code. The move is to shift the centres from production to engineering: invest in senior review capability offshore and nearshore, not only onshore, so the bottleneck does not land on a handful of client-facing leads. Then price on outcomes rather than hours, before your clients ask you to.

Budget. AI tooling is not just a licence per seat. Agentic tools consume compute per task, and that cost scales with usage, not headcount. Treat it as an operating cost to manage, with the same scrutiny as cloud spend.

What technology leaders should do now

  1. Measure outcomes, not output. Track lead time, change failure rate, recovery time and customer-facing delivery. Lines of code and pull request counts will go up regardless, and tell you nothing.
  2. Fix the pipeline before you scale the tools. Invest in automated testing, CI/CD and observability first. That is what turns extra code into extra value instead of extra incidents.
  3. Protect your reviewers. Set limits on pull request size, make agent-generated changes visible, and give senior engineers time for review as real work, not overhead.
  4. Invest in boundaries. Clear domains, documented API's and modular architecture make agents useful. A messy codebase makes them dangerous.
  5. Decide where the lights may go out. Choose deliberately which systems agents may build and review on their own, and keep human understanding of everything else. Treat specs and tests as first-class code.
  6. Keep the junior pipeline open. Hire fewer if you must, but do not stop. Redesign their learning around judging code, not just producing it.
  7. Revisit your sourcing model. Whether you buy capacity or run your own centres, move from pricing hours to pricing outcomes, quality and accountability.
  8. Run controlled experiments. Compare teams or periods with and without specific tools on the same delivery metrics. Your own data beats any vendor benchmark, including the ones in this article.

Conclusion

AI coding agents do not fix an engineering organisation. They reveal it. Strong delivery foundations turn them into a real advantage; weak ones turn them into faster chaos.

That makes this a leadership question, not a tooling question. The organisations that win will be the ones that measure honestly, invest in the unglamorous plumbing, protect the people who exercise judgement and keep growing the next generation who will need it.