Skip to content

AI-ready data foundations · Production DataOps & pipeline engineering · Fractional data leadership · CI/CD for data pipelines · Schema drift, caught before it ships · RAG-ready semantic layers · Infrastructure as code · North America, remote-first

← All insights

Perspective

The developers who felt faster were slower

What the evidence says about AI and this job.

July 10, 2026 · Perspective · Aeolus Data

A copperplate engraving of a pantograph copying instrument, hatched in deep navy on cream paper, a human hand guiding a stylus at one end while the jointed arms reproduce the line at the other, with the guided stylus tip tinted pink.

Most advice about staying employable as a data engineer is a list of things to learn. Add vector databases, add GPU compute, add MLOps, add whichever framework shipped last quarter. The lists are not wrong so much as untested: they assume the effect of AI on this work is known, when the most careful measurement of it produced a result almost nobody expected.

The study worth knowing

In July 2025, METR published a randomised controlled trial on AI-assisted development. Sixteen experienced open-source developers, averaging five years on mature repositories with more than 22,000 stars and over a million lines of code, worked through 246 real issues from their own projects, randomly assigned to allow or disallow AI tooling.

Developers were 19 percent slower when allowed to use AI.

The part that should give everyone pause is not the slowdown. It is the perception gap. Before starting, participants expected a 24 percent speedup. Afterwards, having actually been slower, they still believed they had been roughly 20 percent faster.

Expected, believed, and measured effect of AI assistance on task time

Positive values are speedup. The measured result is a slowdown.

Item Effect on task completion time
Expected beforehand +24%
Believed afterwards +20%
Actually measured slower, not faster −19%
METR, 'Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity', 10 July 2025. Randomised controlled trial, 16 developers, 246 real issues.

METR is careful about scope and so are we. This is sixteen developers, on codebases they already knew deeply, using early-2025 tooling. It does not show that AI assistance fails generally, and METR says so explicitly. What it does show is that self-reported productivity is unreliable evidence, which matters because self-reported productivity is what almost every other claim in this area rests on.

An earlier study by Uplevel, tracking roughly 800 developers over three-month windows, similarly found no significant improvement in pull request cycle time or throughput, and 41 percent more bugs among Copilot users. Set that against GitHub’s own claim of a 55 percent speedup and note who ran which study.

What practitioners actually report

The 2025 Stack Overflow Developer Survey found usage climbing to 84 percent while trust fell to 29 percent. The detail underneath is sharper: 66 percent said AI output is “close but misses the mark,” and 45 percent said debugging AI-generated code takes longer than writing it themselves. Experienced developers were the most sceptical group.

That is a coherent picture rather than a contradictory one. The tools are genuinely useful for a class of work and genuinely costly for another, and the cost lands on exactly the tasks that distinguish a senior engineer: understanding an unfamiliar system, deciding what should exist, and noticing that a plausible answer is wrong.

The labour market is worse, but not in the way you have heard

Two things are true and they get conflated constantly.

Stanford’s Digital Economy Lab, using ADP payroll data across the largest US payroll provider, published Canaries in the Coal Mine in August 2025. It found workers aged 22 to 25 in AI-exposed occupations saw roughly a 16 percent relative employment decline, controlling for firm-level shocks, while experienced workers in the same firms held steady. The effect showed up through hiring rather than wages, and concentrated where AI automates rather than augments.

That is real, rigorous, and concerning. But the viral version, a “67 percent collapse in junior hiring,” has no traceable methodology behind it. It appears to be an inflation of Stanford’s figures by content sites, and it should not be repeated.

Meanwhile Indeed Hiring Lab’s November 2025 report complicates the picture usefully. Data and analytics had the lowest job postings index of any sector it tracks, with rising applications per posting, so the market is genuinely oversupplied. But entry-level roles “remained steady, or even increased, as a share of total postings.” The contraction is broad, not specifically an entry-level purge.

Both things hold: fewer roles overall, and juniors in AI-exposed work hit hardest in payroll data, without entry-level positions disappearing from postings as a share.

What actually appreciates

Here we have to be candid that the evidence is thinner. The labour market and productivity research is solid; the “skills of the future” material is mostly vendor blogs and opinion, ours included. Treat the following as a reasoned position rather than a finding.

The tasks AI handles well are the ones with abundant public examples and a checkable answer: boilerplate transformations, first-draft documentation, test scaffolding, routine bug fixes. The tasks it handles poorly are the ones requiring context that exists nowhere in writing. Nobody has published your company’s actual definition of an active customer, the reason a table has two timestamp columns, or which upstream team will silently change a schema.

That points somewhere specific. Data modelling appreciates, because deciding what should exist is not a retrieval problem. Verification appreciates, because a machine producing plausible output at speed creates more checking work, not less. And writing things down appreciates, oddly, since a model that reads your documentation is only as good as the documentation.

The pantograph is the right image for this. It reproduces a line faithfully at a different scale and produces nothing at all without a hand guiding the stylus.

What we would tell a junior

The honest version is uncomfortable: the entry path is harder than it was, and the tasks that used to build competence are the ones most easily generated.

If you generate the boilerplate you never learned to write, you arrive at the senior conversation without the intuition that used to come free. So do the first pass yourself and use the model to critique it, rather than the reverse. That is slower and it is how the judgment gets built.

Then go deep rather than wide. One cloud, one orchestrator, one transformation layer, held long enough to have been on call for them. The lists of twelve technologies to learn are, if anything, actively harmful advice in a market where the differentiator is depth of understanding rather than breadth of exposure.

Aeolus view. We are a two-founder consultancy, so our bias is obvious: we sell senior judgment. But the market signal points the same way, and we would rather say the uncomfortable half. The work that is genuinely at risk is the work that could be specified precisely enough to be handed to a contractor sight unseen, because that is the same work you can hand to a model. What is not at risk is being the person who can tell that the answer is wrong, and that capability is built by doing the work, which is exactly what the tools now offer to skip.

Where this leaves you

The single most useful thing in the research is not any figure about jobs. It is that the developers in the METR trial were wrong about their own productivity, in the favourable direction, after the fact. If experienced engineers cannot accurately assess whether a tool made them faster on their own codebase, then almost every confident claim in this discussion, including the optimistic ones, is built on the same unreliable instrument.

Measure. Even crudely. Cycle time on a class of tickets before and after, or the share of generated code that survives review unchanged. You will learn more from a month of that than from any amount of commentary, including this.

If you are trying to work out which capabilities your data team should be building, we are happy to think it through with you. Usually the answer involves fewer new tools than expected.

Want a second opinion on your data stack?

Every Aeolus engagement starts with a fixed-fee data & AI-readiness audit — a short, low-risk first step before any larger build.

Book a data & AI-readiness audit