Skip to content

AI-ready data foundations · Production DataOps & pipeline engineering · Fractional data leadership · CI/CD for data pipelines · Schema drift, caught before it ships · RAG-ready semantic layers · Infrastructure as code · North America, remote-first

← All insights

Perspective

Why AI-Assisted Developers Feel Faster Than They Are

New research challenges claims of AI speedups.

July 10, 2026 · Perspective · Leon Liang

A copperplate engraving of a pantograph copying instrument, hatched in deep navy on cream paper, a human hand guiding a stylus at one end while the jointed arms reproduce the line at the other, with the guided stylus tip tinted pink.

Most advice for data engineers focuses on a growing checklist of tools: vector databases, GPU compute, MLOps, or whatever framework launched last quarter. These lists aren’t necessarily wrong, but they are untested. They assume the impact of AI on the profession is a settled fact, when the most rigorous measurements suggest something entirely different.

The study worth knowing

In July 2025, METR published a randomised controlled trial on AI-assisted development. The study followed sixteen experienced open-source developers working on mature repositories (22,000+ stars, 1M+ lines of code). They tackled 246 real-world issues, randomly assigned to either use or avoid AI tooling.

The result: Developers were 19 percent slower when using AI.

The most striking finding wasn’t the slowdown, but the perception gap. Before the trial, participants expected a 24 percent speedup. After the trial—despite actually being slower—they still believed they had been roughly 20 percent faster.

Expected, believed, and measured effect of AI assistance on task time

Positive values are speedup. The measured result is a slowdown.

Item Effect on task completion time
Expected beforehand +24%
Believed afterwards +20%
Actually measured slower, not faster −19%
METR, 'Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity', 10 July 2025. Randomised controlled trial, 16 developers, 246 real issues.

METR is careful about scope, and we should be too. This study involved a small group of developers working on familiar codebases with early-2025 tools. It doesn’t prove AI assistance fails across the board. However, it does prove that self-reported productivity is an unreliable metric—which is problematic, because almost every claim about AI productivity relies on it.

This aligns with an earlier study by Uplevel. Tracking roughly 800 developers over three months, Uplevel found no significant improvement in throughput or pull request cycle time, while noting 41 percent more bugs among Copilot users. Contrast this with GitHub’s own claim of a 55 percent speedup, and the source of the data becomes critical.

What practitioners actually report

The 2025 Stack Overflow Developer Survey reveals a tension: usage has climbed to 84 percent, but trust has fallen to 29 percent.

The details are telling:

  • 66 percent say AI output is “close but misses the mark.”
  • 45 percent say debugging AI-generated code takes longer than writing it from scratch.
  • Experienced developers are the most sceptical.

This isn’t a contradiction; it’s a pattern. AI is highly efficient for certain tasks and costly for others. The cost is highest on the exact tasks that define a senior engineer: navigating unfamiliar systems, architectural decision-making, and spotting plausible but incorrect answers.

The labour market: Signal vs. Noise

Two distinct trends are often conflated in the current discourse.

First, the rigorous data. In August 2025, Stanford’s Digital Economy Lab published Canaries in the Coal Mine, using ADP payroll data. They found that workers aged 22–25 in AI-exposed occupations saw a roughly 16 percent relative employment decline (controlling for firm-level shocks), while experienced workers remained steady. This effect manifested in hiring rather than wages, concentrated where AI automates rather than augments.

Second, the viral myth. Claims of a “67 percent collapse in junior hiring” lack a traceable methodology. This appears to be an inflation of the Stanford figures by content sites and should be ignored.

Adding further nuance, the Indeed Hiring Lab’s November 2025 report shows that data and analytics have the lowest job postings index of any tracked sector, with rising applications per posting. The market is oversupplied. However, entry-level roles “remained steady, or even increased, as a share of total postings.”

The reality: The contraction is broad. While juniors in AI-exposed roles are hit hardest in payroll data, entry-level positions aren’t disappearing as a percentage of overall postings.

What actually appreciates in value

The evidence for “skills of the future” is thinner than the labour data, often consisting of vendor blogs and opinion. Consider the following a reasoned position rather than a proven finding.

AI excels at tasks with abundant public examples and checkable answers: boilerplate transformations, first-draft documentation, test scaffolding, and routine bug fixes. It fails at tasks requiring context that isn’t written down. A model doesn’t know your company’s specific definition of an “active customer,” why a certain table has two timestamp columns, or which upstream team is likely to change a schema without notice.

This creates three areas of appreciation:

  1. Data Modelling: Deciding what should exist is a design problem, not a retrieval problem.
  2. Verification: As machines produce plausible output at speed, the demand for rigorous checking increases.
  3. Documentation: A model is only as good as the context it reads. Clear, written specifications become more valuable, not less.

The pantograph is the perfect metaphor here: it reproduces a line faithfully at a different scale, but it produces nothing without a hand guiding the stylus.

Advice for the junior engineer

The uncomfortable truth is that the entry path is harder. The “grunt work” that used to build foundational competence is now the easiest work to automate.

If you rely on AI to generate the boilerplate you never learned to write, you will reach senior-level conversations without the intuition that used to come for free. To avoid this, do the first pass yourself. Use the model to critique your work, rather than using the model to do the work. It is slower, but that is how judgment is built.

Focus on depth over breadth. Master one cloud, one orchestrator, and one transformation layer—and stay with them long enough to be on call for them. In a market where the differentiator is depth of understanding, a checklist of twelve different technologies is actively harmful advice.

Aeolus view — As a two-founder consultancy, we sell senior judgment. Our bias is clear, but the market signal supports us. The work genuinely at risk is work that can be specified precisely enough to be handed to a contractor sight unseen—because that is the same work you can hand to a model. What remains safe is the ability to tell when an answer is wrong. That capability is built by doing the work, which is exactly what the tools now offer to skip.

Where this leaves you

The most important takeaway from the METR research is that experienced engineers were wrong about their own productivity. If experts cannot accurately sense whether a tool is making them faster on their own codebase, then most confident claims in this debate are built on an unreliable instrument.

Measure your own progress, even crudely. Track cycle times on specific ticket classes before and after adopting a tool, or measure the percentage of generated code that survives review unchanged. A month of your own data is worth more than any amount of commentary.

If you are determining which capabilities your data team should prioritize, we can help you think it through. If your data isn’t ready for this yet, we’ll help you get there. Usually, the answer involves fewer new tools than you’d expect.

Want a second opinion on your data stack?

Every Aeolus engagement starts with a fixed-fee data & AI-readiness audit — a short, low-risk first step before any larger build.

Book a data & AI-readiness audit