Junior engineers and AI is a measurement problem, because speed with the tool tells a leader little about understanding without it. Garima Agarwal, a back-end engineer and tech leader, makes that case in a recent article, and this piece extends one claim from it.
Agarwal’s team shipped 14 features and built no mental models
Agarwal describes her own team after six months of mandatory AI coding tool use. The team shipped 14 features, and the juniors had built no working picture of the systems behind them. The failure became visible on a production incident. A race condition appeared in an order processing service, and the junior engineer on call had built a feature on that same service two months earlier, yet could not work through the problem.
Her response was to restructure onboarding and leave the tools in place. She separates output tasks, where AI is encouraged, from learning tasks, where the engineer attempts the work first. She added weekly 15-minute decision walks in which a junior explains the reasoning behind a merged pull request, a two-week codebase read at the start of onboarding, and monthly debugging sessions built on anonymized production incidents. Her own measure of progress is time to meaningful contribution, defined as the first incident a junior can work through independently, which fell from seven months to four. In her words, "Speed helps you this quarter. Understanding helps you for years."
Performance with the tool does not predict performance without it
Two studies bear directly on her claim. In a field experiment published in the Proceedings of the National Academy of Sciences in 2025, Hamsa Bastani and colleagues gave nearly 1,000 high school math students access to a generative AI tutor during practice. Students with a standard ChatGPT-style interface raised their practice grades by 48%. When access was removed, those students performed 17% worse than students who never had access. A second tutor, built with prompts designed to protect learning, largely prevented that harm.
Anthropic’s randomized trial on coding skills, published on 29 January 2026, found the same shape in software. Fifty-two mostly junior engineers learned an unfamiliar Python library, Trio, either with an AI assistant or by hand. The AI group scored 50% on a follow-up quiz and the hand-coding group scored 67%. The AI group finished about two minutes faster, and Anthropic reports that time difference was not statistically significant. Participants who asked conceptual questions, or who asked for explanations alongside generated code, scored 65% or higher, although those groups were small, with between two and seven participants each.
Both studies measured understanding in a condition where the tool was absent. That is the feature a leader needs to copy. Output measured with the tool present can rise while understanding falls, and the two studies show both movements in the same participants.
An unassisted checkpoint turns that evidence into a management measure
Agarwal’s 1-to-1 test fits this design. She asks an engineer to explain how a service they shipped actually works. That question removes the tool and the pull request from the room and leaves the engineer’s own model. A leader can repeat it on a fixed cadence for every junior engineer and record only whether the explanation held up, with no score attached. The decision walk does the same job for a single change, and the first independent incident does it for a live system.
The Anthropic result adds a design hint. Because the learning-preserving patterns involved asking for concepts and explanations, a team policy of "explain before you merge" fits how the better performers already worked. The PNAS result adds another, since the safeguarded tutor largely avoided the harm. Tool configuration is a lever a leader controls, and restricting AI is only one option among several.
The boundary conditions matter. Agarwal’s numbers come from one team and are self-reported, with no comparison group, so the seven-to-four-month improvement cannot be separated from the effect of simply paying attention to juniors. Her incident-based measure also depends on incidents occurring, and a reliable system produces few, so the signal arrives slowly exactly where operations are healthiest. Her article does not say whether AI tools are allowed during those incidents, and in real work an engineer will often have them open. The Anthropic trial used a short task with an unfamiliar library and a quiz, and the PNAS study used high school students, so neither is a direct measurement of junior engineers on a production codebase. What they support is the direction of the effect and the value of an unassisted check. They do not support an exact size for it.
The cost side is small by comparison. Agarwal reports that the decision walks take about 30 minutes of senior engineer time per week and that the codebase read adds two weeks to onboarding. Those figures are hers, and a different team should expect different numbers.
Why This Matters for Engineering Leadership
Leaders usually see junior engineers through output: pull requests merged, features shipped, tickets closed. All of those rise with AI assistance, so they cannot show whether anyone understands the system. A standing unassisted checkpoint gives leaders a second reading that the tool cannot inflate. From an AI governance view, the same logic applies to accountability. A person who ships a change must be able to explain it, and a junior engineer who cannot is not yet an accountable owner, however fast the change merged. Teams that build this check into onboarding early will find out about a comprehension gap in a one-to-one conversation, where it is cheap, and not during an incident.
FAQ
Why can AI make junior engineers faster without making them better?
Output measured with the tool present can rise while understanding falls. In a PNAS field study, students with AI access improved during practice and then performed 17% worse than peers once access was removed.
What did Anthropic’s 2026 coding skills trial find?
In a randomized trial of 52 mostly junior engineers learning the Trio library, the AI-assisted group scored 50% on a quiz and the hand-coding group scored 67%. The two-minute speed difference was not statistically significant.
What is time to meaningful contribution?
Agarwal defines it as the first incident a junior engineer can work through independently. She reports it fell from seven months to four after her team restructured onboarding.
What is a decision walk?
It is a weekly 15-minute session in which a junior engineer explains the reasoning behind a merged pull request. Agarwal reports it costs about 30 minutes of senior engineer time per week.
Does Agarwal recommend restricting AI tools for juniors?
No. She recommends restructuring onboarding so output tasks and learning tasks are separate, and she keeps AI in use for output tasks.
How can a leader test comprehension without a score?
Ask the engineer to explain how a service they shipped works, without the tool, and record whether the explanation held up. This is a pass-or-hold judgment, and the checkpoint design is this piece’s proposal.
