In March 2026, Andrew Filev published a number that stopped engineering leaders cold: his team of 30 engineers at Zencoder was producing 170% of the output of the previous 36-person team. Not 20% more. Not 50% more. One hundred and seventy percent — at 80% of the headcount. Verified via six months of JIRA data. And Filev is not a startup founder with something to prove: he previously built Wrike, sold for $2.25 billion. This was not a hot take. This was a case study.
What the Data Actually Shows
The Zencoder result sits at the extreme end of a broader productivity wave that is now measurable at scale. Faros AI’s dataset of 10,000 developers across 1,255 teams shows AI-assisted engineers completing 21% more tasks and merging 98% more pull requests per month. GitHub Copilot, now with 20 million users and 4.7 million paid subscribers (75% YoY growth), generates 46% of all code written by its users on average — reaching 61% for Java developers. McKinsey’s State of AI 2025 puts AI-assisted developers at 40–55% more code per week, with $2.6–4.4 trillion in annual economic value potential from AI in software engineering.
By late 2025, AI-authored code made up 26.9% of all production code. 84% of developers report using AI tools regularly. The throughput story is real — and getting louder.
The Paradox Inside the Paradox
Here is what the productivity narrative is not telling you. Faros AI’s 2026 Engineering Report — drawing on 22,000 developers across 4,000 teams — shows a picture that is far more complicated than the headline numbers suggest:
- Epic completion up 66.2%, task throughput up 33.7%
- Bug rate per PR up 28%
- Code churn up 861%
- Production incidents per PR tripled
- Code review time extended 5x
- 31.3% more PRs merged without review
The throughput is real. The quality problem is also real. AI is accelerating code generation faster than engineering processes can absorb it — and the bottleneck has simply moved downstream, from writing to reviewing to debugging in production. Output measured in story points looks great. Output measured in production incidents looks alarming.
This is not an argument against AI tools. It is an argument that teams measuring only throughput are managing half the picture. DORA metrics — deployment frequency, lead time, change failure rate, time to restore — remain stubbornly flat even as PR velocity climbs. The engineering process is being stress-tested in ways that velocity metrics cannot capture.
The Org Chart Is Breaking
The deeper shift is structural, not tooling-level. InteligenAI’s 2026 analysis describes what is emerging as “Atomic Teams” — small, vertically integrated pods where one engineer with AI agents can own an entire product slice from architecture to deployment. The traditional horizontal engineering org (frontend team, backend team, QA team, DevOps team) is being replaced by vertical ownership at the pod level.
Filev’s framework — which he calls the “double funnel” — describes the geometry precisely. The old sprint model was a diamond: a small group defining requirements, expanding to a large engineering team, narrowing again at QA. The new model inverts this: deep human engagement at the beginning (architecture, problem framing, acceptance criteria), AI execution in the middle, and deep human engagement again at the end (review, validation, deployment decisions). Sprint cadence compresses from two weeks to 36-hour micro-cycles.
McKinsey’s State of Organizations 2026 confirms the workforce implications are already being felt: 30% of organizations expect AI-driven workforce decreases, and 23% are already scaling AI agent systems. The historical parallel is instructive: when the spreadsheet arrived in 1979, 400,000 accounting clerk jobs disappeared — and 600,000 accountant jobs were created. The skills mix shifted upward, not away.
What Engineering Leaders Should Do Now
Three decisions matter more than any specific tool choice in 2026. First: measure both directions. If your engineering metrics only track throughput and velocity, you are flying with one instrument. Add bug rate per PR, code churn, and production incident rate to your dashboard before celebrating the numbers. Second: redesign review, not just writing. The 5x extension in review time is not a problem to solve with more reviewers — it is a signal that review processes were designed for human-generated code and need to be rebuilt for AI-generated code at higher volume. Third: plan for the org chart change deliberately. Atomic teams do not emerge from policy memos. They emerge from engineering leaders who intentionally restructure ownership, accountability, and on-call rotations around the new unit of work.
Filev’s 170% result is the best-documented proof that the upside is real and achievable. The Faros AI data is the best-documented proof that capturing the upside without managing the downside creates a different kind of problem. Both are true simultaneously — and the engineering leaders who hold both truths in their heads at once are the ones who will figure out what comes next.
