Mover

Cognitive debt in AI coding: the evidence

Updated .

Cognitive debt is the gap between what your code does and what you understand about it. Researcher Margaret-Anne Storey argues AI agents make it grow faster. In Anthropic's trial, developers who used AI to learn a new library scored 17% lower on a quiz about it, and the biggest gap was debugging.

Last verified 30 September 2026. Every study and quote below links its source.

Where does the term come from?

Margaret-Anne Storey, who studies software teams, describes it as debt that lives in people: "the humans involved may have simply lost the plot and may not understand what the program is supposed to do, how their intentions were implemented, or how to possibly change it."DX newsletter Her example is a student team that, by week seven or eight, could not make simple changes without breaking something, because nobody could explain why the design was the way it was. Simon Willison has written about the same thing in his own projects, getting lost in code he had prompted into existence without reviewing it.Willison

What did the Anthropic trial measure?

52 mostly junior developers, all new to Trio, an async Python library, built two features in 35 minutes. Half had an AI assistant that could see their code. Afterwards everyone took a quiz without AI. The AI group averaged 50% and the hand coding group 67%, a gap the researchers call nearly two letter grades. About two minutes faster was all the AI group gained, and that gap was not statistically significant.Anthropic

The six ways people used the assistant split cleanly. Full delegation, gradual handover and asking the AI to fix each error scored 24% to 39%. Asking only conceptual questions, asking for explanations with the code, or generating code and then questioning it scored 65% to 86%. The quiz came right after the task, so it says nothing about memory a month later, and some pattern groups had only two people.arXiv

Which habits keep a codebase yours?

  1. Write the intent before the agent writes code. Two or three lines on what the change should do and what it must not touch. Put lasting decisions in your CLAUDE.md or AGENTS.md so the next session inherits them.
  2. Ask for the explanation with the diff. This was one of the three high scoring patterns in the trial.
  3. Quiz yourself after the agent works alone. Three questions on what changed, answered before you read the diff again. If you cannot answer one, read that part.
  4. Debug something by hand every week. Debugging showed the largest gap in the trial, so it is the skill to practise on purpose.
  5. Keep a short decision log. One line per choice and the reason. Storey's warning signs include knowledge held by only one or two people and a system that starts to feel like a black box.DX newsletter

When can the agent carry it?

Simon Willison makes the point that for simple code, such as fetching some data and returning it as JSON, you can try the feature, guess how it works and glance at the code to be sure. The debt matters when the core of the application becomes something you cannot reason about.Willison Anthropic's write up also points to its earlier observational research, which found large time savings on tasks where people already had the skill.Anthropic

Where does Mover OS fit?

Mover OS runs inside Claude Code, Codex or Gemini CLI. Two parts of it touch this problem. After a session where the agent made five or more commits or large changes you did not watch, its log workflow offers a three question quiz on what changed and reveals the answers with file and line after you try. And work the agent did stays marked unverified in your plan until you check it and say it is done.

The quiz is optional and only offered after heavy autonomous sessions. Mover does not stop an agent writing a large diff, and it has not been tested for its effect on anyone's understanding. Mover OS is not for anyone who wants a phone app or a chat window: it needs one of those three terminal agents.

Common questions

Is cognitive debt the same as technical debt?

No. Technical debt lives in the code and usually shows up as friction when you change it. Cognitive debt, in Storey's framing, lives in the people and shows up as not knowing why the code is the way it is.

Did AI make the developers in the Anthropic study faster?

On average about two minutes faster over a 35 minute task, and that difference was not statistically significant. People who fully delegated were faster but learned the least.

Does Claude Code have a mode for learning?

Anthropic's write up of the trial names Claude Code's Learning and Explanatory modes as options designed to build understanding. The trial did not test them.

Sources

  1. Anthropic 2026, How AI assistance impacts the formation of coding skills
  2. How AI Impacts Skill Formation (arXiv 2601.20245)
  3. Margaret-Anne Storey on cognitive debt, DX newsletter, 22 April 2026
  4. Simon Willison, posts tagged cognitive debt

Mover OS: $49 once. Runs inside Claude Code, Codex or Gemini CLI. Offers a quiz on what the agent finished while you were away. Does not run while the agent is closed.