The Code Is Cheap. Understanding Is Not.
Why AI Coding Needs Human Understanding

The Code Is Cheap. Understanding Is Not.
AI agents are making software cheap to produce, but Geoffrey Litt argues that the scarce resource in an agent-driven industry will be the engineer who still understands what the system does, why it was built that way and how to change it next. His warning is not that machines cannot write code; it is that teams may mistake a working output for a working mental model.
Litt, a design engineer at Notion and a former researcher at Ink & Switch, made that case at the AI Engineer conference in July 2026. He described printing AI-generated explanations of his own code changes, taking them to a coffee shop and reading them like a textbook. The joke landed because it sounded excessive. The argument did not: when software can be generated faster than people can absorb it, understanding becomes the limiting factor.
The Review Queue Is Not the Real Problem
The fashionable description of software development says that code review is the last human bottleneck. Agents can write functions, modify files, run tests and propose pull requests. Humans remain in the process, according to this account, because they must inspect the work, catch errors and approve the result. Once verification becomes sufficiently reliable, the human reviewer can be removed or reduced to a final exception handler.
That account is not entirely wrong. A system that produces incorrect software still needs supervision. Tests fail, specifications are ambiguous and an agent can satisfy the literal wording of a request while missing the purpose behind it. A competent reviewer has to notice those failures. But verification is only one reason to read a change, and it is not the reason that matters most to the future of engineering.
Litt’s sharper claim is that close review changes the person doing it. The engineer does not merely decide whether a patch is safe. He acquires a more detailed model of the system: which assumptions it makes, where data moves, which constraints are accidental and which are essential. That model becomes material for the next decision. It determines what the engineer can ask an agent to do, what trade-offs can be considered and which ideas are even visible.
To treat review as a vote for or against correctness is to reduce a creative activity to quality control. A reviewer who has learned nothing from a change may approve it, but remains no better equipped to direct the project. The code may be correct today while the team becomes less capable of deciding what the code should become tomorrow.
That distinction matters because agents are not replacing a single clerical task. They are compressing a sequence of tasks that once forced engineers to build familiarity: searching through a repository, tracing a dependency, testing a hypothesis, writing an awkward first version and discovering where the abstraction breaks. If all of those experiences are delegated, the team can gain output while losing the knowledge that gives output direction.
The result is a peculiar inversion. The more capable the agent becomes at checking its own work, the less useful it is to define the human role as checking. Human participation has to move up a level. The question is no longer only whether the implementation passes. It is whether anyone on the team understands the implementation well enough to generate the next good question.
The Productivity Warning Behind the Promise of Speed
The promise of AI coding is simple: give a developer a machine that can produce software and the developer will finish more work in less time. Yet the early evidence was less tidy than the sales pitch. A 2025 randomised controlled trial by METR studied 16 experienced open-source developers working on projects they knew well. Across 246 real issues, the developers who were allowed to use early-2025 AI tools took 19 per cent longer to complete the work than those who worked without them.2
That finding has since been overtaken by newer evidence and should not be treated as a permanent verdict on AI-assisted programming. METR itself describes the result as a snapshot of an earlier generation of tools, and later data points in a different direction. But the original experiment exposed a problem that remains even when the models improve: generating a plausible change is not the same as integrating it into a living system.
The developers in the trial were not novices who had never used an assistant. They worked in large, mature repositories with high standards and implicit requirements that were not always captured by an automated test. They had to understand local conventions, preserve behaviour, maintain documentation and produce a change that another maintainer would accept. In such environments, the cost of unfamiliar output can exceed the time saved by producing the first draft.
The experiment also exposed a more uncomfortable gap between perception and performance. Before using the tools, developers expected AI to reduce completion time by 24 per cent. After the work, they still believed the tools had made them 20 per cent faster, even though the measured result showed a slowdown. The lesson is not that developers are foolish. It is that fluency is a poor substitute for measurement. A stream of generated code feels like progress even when the surrounding work has become harder to understand.
That feeling is likely to become more powerful as agents operate for longer periods and produce larger changes. A one-line suggestion can be inspected at a glance. A multi-file refactor, migration or feature branch cannot. The agent may have completed its instructions, but the human inherits a new task: reconstructing the reasoning that led to the result and finding the parts the tests do not describe.
This is why the bottleneck argument should not be reduced to a race between faster models and slower people. The relevant measure is not how quickly a machine can emit code. It is how quickly a team can turn generated code into reliable knowledge, shared judgement and a clear next move. If the answer is “not quickly enough”, the organisation has not automated the bottleneck. It has moved the bottleneck into the heads of its staff.
Correctness Cannot Tell a Team What to Build
Verification is powerful because it answers a bounded question. Does the program satisfy this test? Does the database migration complete? Does the interface respond to the expected input? Has the agent followed the stated instruction? Those questions are necessary, but they are not sufficient for serious engineering.
Software projects are full of decisions that cannot be settled by a green test suite. A feature can be implemented correctly and still be the wrong feature. A data structure can meet the current requirements while making future changes needlessly expensive. A user interface can behave exactly as specified while encouraging the wrong behaviour. A service can remain within its performance budget and yet become impossible for the next engineer to maintain.
Specifications are not neutral containers of intent. They are partial descriptions written by people who do not yet know everything that will matter. The act of building often reveals missing requirements, contradictions and opportunities. Engineers who understand the system can recognise those discoveries as they happen. Engineers who are separated from the implementation by an agent may have difficulty distinguishing a genuine improvement from an attractive accident.
This is the point at which the human role becomes creative rather than supervisory. Understanding gives a person a vocabulary for manipulating a system. Without that vocabulary, prompting becomes a form of guessing. The user can request “make this faster”, “simplify the architecture” or “add a better workflow”, but cannot judge which internal changes would achieve those goals without causing damage. The agent can propose options, yet the quality of the conversation depends on the human having concepts to recombine.
Litt’s phrase is blunt: understanding is the foundation for having the next idea. That is not a mystical claim about human superiority. It is a description of how expertise works. People who know the components of a system can see relationships that are invisible to someone who sees only its surface. They can notice that a performance problem is really a data-model problem, that an interface complaint reflects a policy choice or that a proposed feature would eliminate a useful constraint.
There is also a distinction between knowing that a system works and knowing how it works. The first supports a demonstration. The second supports maintenance, explanation and invention. Agents are good at producing demonstrations. Teams need the second kind of knowledge if they are to remain owners of their work.
That makes the popular dream of a fully autonomous coding loop less attractive than it first appears. A loop that writes, tests and merges without human intervention may produce a large volume of software. It does not automatically produce a group of people capable of explaining the system, challenging its assumptions or deciding where the project should go. Autonomy can remove human labour from the loop. It cannot remove the need for human purpose.
Cognitive Debt Is the Bill for Moving Too Fast
The familiar engineering warning is technical debt: shortcuts that make a system harder to change later. AI introduces a related liability. Cognitive debt accumulates when a person accepts changes without developing a reliable understanding of the system those changes create. The code may be clean. The tests may pass. The debt still exists in the gap between the software and the team’s mental model.
Researcher Margaret-Anne Storey describes cognitive debt as the burden created when people move quickly without preserving the shared theory of how a system works. Simon Willison, a developer and writer, has given the idea a wider audience. He has described prompting entire features into existence without reviewing their implementations, finding that the method worked until he became lost in his own projects and could no longer make confident decisions about what to add next.3
The phrase is useful because it shifts attention from the visible artefact to the people who must live with it. Technical debt can often be located in a file, a dependency or a database table. Cognitive debt is harder to see. It appears as hesitation in planning meetings, repeated rediscovery of the same behaviour, an inability to estimate a change or a growing dependence on the agent that created the original system.
It also compounds. A developer who does not understand one feature can ask an agent to modify it. The new change introduces another layer of unfamiliarity. The developer then asks for a third change without resolving the first two. Each individual step looks manageable, but the accumulated system becomes opaque. At some point, the person is no longer directing a tool. He is negotiating with a machine that has a more complete map of the project than he does.
This is not the same as saying that every line of code must be memorised. No serious engineer carries an entire codebase in his head. Understanding is selective and structural. It means knowing the central abstractions, the boundaries between components, the reasons behind unusual decisions and the failure modes that matter. It means being able to investigate the details when necessary because the larger map is intact.
Teams can incur cognitive debt collectively as well as individually. If each engineer uses a private agent conversation to generate private solutions, the organisation acquires a set of incompatible explanations. One person understands the repository as a set of services; another understands it as a set of workflows; a third knows only the prompts that seem to make the tests pass. When the original author leaves, the system loses not just a programmer but an interpreter.
The cost eventually reaches management. Estimates become less reliable because nobody knows which hidden dependencies a change will disturb. Hiring becomes more difficult because new staff need to learn a system that has not been explained. Incidents take longer to resolve because the team cannot identify which behaviour is deliberate. The firm may still boast about its output, but the output is being purchased with future decision-making capacity.
Turn Every Significant Change Into a Lesson
Litt’s first remedy is not a ban on agents. It is a better form of explanation. A raw diff records what lines changed, but it rarely explains the world in which those lines make sense. It may show that a condition was moved, a function was split or a file was added. It does not necessarily show the problem the change addresses, the assumptions it relies on or the alternatives that were rejected.
His explain-diff workflow asks the agent to produce a teaching document before the change is treated as understood. The document begins with the background: how the system is organised, which coordinate system or data model is relevant and which subsystems are involved. It then states the intuition behind the change in plain language. Only after that does it walk through the code.
The order matters. Engineers often approach a difficult diff in the order the tool presents it, which is rarely the order a human needs. One file may be changed first because it was the easiest place to start, while another file contains the conceptual key. A teaching document can reverse the sequence. It can introduce the architecture, explain the causal chain and present the implementation as evidence for an idea rather than as an unexplained pile of edits.
Litt calls this a literate diff: prose and code arranged to make the change learnable. The format is closer to a short textbook chapter than to a pull-request summary. It can include diagrams, examples and interactive figures where those reduce the burden on the reader. The point is not ornamental documentation. The point is to make the reasoning recoverable by someone who did not watch the agent work.
There is a practical advantage here for teams that are worried about review time. A good explainer does not eliminate inspection, but it moves inspection from blind reconstruction to informed questioning. The reviewer can ask whether the explanation matches the intended design, whether a hidden assumption is justified and whether the implementation actually follows the stated logic. Without an explanation, the reviewer must infer the entire argument from the final code.
The document also creates an artefact that can outlast the original interaction. Agent conversations are often transient and difficult to search. A durable explanation can sit with the design notes, be commented on by colleagues and be revised when the system changes. It gives the team a shared account of what happened rather than leaving the most important reasoning inside a private chat window.
This method exposes an inconvenient truth about automation. The machine can reduce the cost of writing software, but it cannot make comprehension disappear. Someone still has to decide what the change means. The sensible response is to ask the machine to help with that work rather than pretending the work is unnecessary.
The Five-Question Test for False Confidence
Reading an explanation can create its own illusion of mastery. A clear document feels understandable because the sentences are familiar and the diagrams look persuasive. That feeling is not proof. Andy Matuschak’s work on learning makes the distinction explicit: people often confuse exposure to an explanation with the ability to retrieve and use the underlying idea. His essay “Why books don’t work” argues that static prose leaves much of the burden of checking understanding to the reader.4
Litt’s answer is a small quiz at the end of every explainer. He asks the agent to generate five medium-difficulty questions about the change and refuses to send agent-written code to a teammate until he can answer them. The rule is deliberately simple. It turns understanding from a mood into a test.
A useful question is not “What did this function do?” That can be answered by repeating its name. Better questions probe causality and boundaries: What assumption makes this approach safe? What would break if this value were missing? Why does the change belong in this layer rather than another? Which behaviour is preserved by the new code? What evidence would show that the design is wrong?
Those questions do two jobs. They reveal gaps in the reader’s model and expose weaknesses in the agent’s explanation. A fluent but shallow summary may describe the intended behaviour while ignoring an edge case. The quiz forces the explanation to become more precise. If the agent cannot produce questions whose answers are clear, the document is not teaching enough. If the engineer cannot answer them, the review is not complete.
The method resembles Quantum Country, the interactive introduction to quantum computing created by Matuschak and Michael Nielsen. It embeds retrieval prompts into the reading so that learners have to recall ideas rather than passively move through a sequence of pages. The project describes itself as a “mnemonic book” designed to make memory part of the reading experience rather than an afterthought.5
Software teams have usually treated quizzes as something for apprentices or schoolchildren. That is an odd prejudice. Engineers routinely test machines because they distrust appearances. They should test their own understanding for the same reason. A person can read a 2,000-line change, recognise the vocabulary and still be unable to predict what happens when one assumption changes.
The quiz is also a useful speed regulator. The pressure surrounding AI encourages immediate acceptance: the code exists, the demo works and another prompt is waiting. Five questions create a pause between production and commitment. They do not make the process slow for its own sake. They prevent a team from confusing the speed of generation with the speed at which it can responsibly absorb change.
Build a Small World for the Parts You Cannot See
When explanation is not enough, Litt recommends building a microworld: a small, temporary environment in which a difficult system can be explored through action. The idea comes from Seymour Papert, the mathematician and educator who argued that learners understand abstract concepts by making things in a context where the ideas become tangible. Papert’s Logo language and turtle robot let children give instructions, observe the result, diagnose the error and try again. His central contribution was not a colourful robot but a different theory of learning: people become more capable by constructing objects they can inspect and modify.6
The equivalent for software is not a production feature. It is a disposable world designed around one confusion. If an engineer cannot understand how an interpreter resolves a query, an agent can build a debugger that shows the internal state at each step. The learner can move backwards and forwards through the process, change an input and watch the consequence. Instead of reading an abstract account of resolution, the engineer develops a feel for the mechanism.
The difference between a microworld and a conventional example is control. Documentation often shows the path chosen by the author. A playground lets the reader choose a path, make a mistake and observe the system’s response. That experience supplies information that a polished explanation suppresses. It shows not only what the machine does when everything is correct, but also where the model breaks.
Litt applied the same idea to a website migration. Rather than ask an agent to move the site and return a completed result, he asked it to create a game-like interface with the old site on one side and the new site on the other. A “next” control performed one migration step at a time. The interface displayed the commands, the emerging pages and the changing file tree. The migration became something he could inhabit instead of a black-box operation he had to trust.
That is a powerful division of labour. The agent handles the repetitive work of building the simulator, preparing the visualisation and wiring together the controls. The human performs the activity that creates understanding: changing conditions, seeing relationships and repairing errors. The machine does not remove the learning process. It makes a learning environment cheap enough to request whenever a subsystem becomes opaque.
Microworlds are particularly valuable for systems whose important behaviour is invisible in the final interface. Compilers, interpreters, build pipelines, asynchronous queues, permission systems and data migrations may all appear simple from the outside while performing a long sequence of hidden operations. A visual timeline, state inspector or toy implementation can expose those operations without requiring the engineer to navigate the full complexity of the production system.
The result is not always a reusable tool. It may be thrown away once the concept is clear. That is not waste. A disposable environment that prevents a week of confused changes can be more valuable than a permanent feature. The deliverable is not software to ship. It is the mental model the engineer carries back to the real system.
Understanding Has to Be Shared, Not Privatised
The individual engineer is not the only person at risk. A team can lose its shared understanding even when every member is busy learning something in private. The modern agent workflow encourages that fragmentation. One developer opens a terminal agent, another uses an IDE assistant and a product manager consults a separate chat. Each receives a fast answer. The group receives no common account of the reasoning.
That matters because software is a social object. A repository is not maintained by the person who last touched it. It is interpreted by designers, product managers, security staff, operators and future engineers. They need a vocabulary for discussing trade-offs. They need to know which decisions are fixed, which are provisional and which were made only because of a temporary constraint.
Litt’s third technique is the shared space: a visible conversation where agents and humans work alongside one another. Plans, explanations, questions and generated artefacts live where colleagues can inspect them. An agent is not a private oracle delivering an answer to one employee. It becomes a participant in a team’s working record.
This approach also changes accountability. A private agent chat can make a decision appear more certain than it was because nobody else sees the uncertainty, the discarded alternatives or the missing context. In a shared space, colleagues can challenge the premise before the implementation hardens. They can add knowledge that was absent from the prompt. The agent’s work becomes an object of collective reasoning rather than a personal shortcut.
Notion’s own product direction illustrates the commercial version of this idea. In its July 2026 release, the company introduced External Agents, allowing Claude and Cursor to be assigned work from a shared board and mentioned alongside teammates. The release also added interactive HTML blocks that agents could create inside documents, including quizzes, calculators and organisational visualisations.7
There is a clear product pitch here, and organisations should be sceptical of any vendor suggesting that a new interface automatically creates collaboration. A shared window can still contain shallow thinking. But the underlying design principle is sound: explanations and agent actions should be placed where the people responsible for the system can see, question and extend them.
The distinction resembles the difference between a meeting and an archive. A meeting may resolve an issue, but an archive lets later people understand how the decision was reached. Agentic work needs both. The live interaction helps the team move. The shared record prevents movement from becoming institutional amnesia.
For managers, this means measuring more than individual throughput. A developer who produces ten changes in private may be less useful than one who produces six changes that the whole team can understand and maintain. The relevant unit is not the number of prompts completed. It is the amount of durable, transferable capability created by the work.
The Forgotten Blueprint Was Drawn in 1972
Litt’s argument reaches back beyond current tools to a question posed by Alan Kay more than half a century ago. In his 1972 paper “A Personal Computer for Children of All Ages”, Kay imagined a computer that would let children build, simulate and alter the worlds around them. The famous image associated with the vision shows children using a screen, but the point was not passive consumption. The children were meant to change the rules of the game itself.8
That distinction has been blurred by decades of computing built around consumption. The dominant device is a portal to finished services: watch, scroll, search, buy and communicate. Creation is available, but it is often treated as a specialised activity reserved for programmers and designers. The computer becomes an appliance rather than a medium for thought.
AI agents could push computing further in either direction. They can make the device an even more efficient consumption engine, predicting what a person wants and removing the need to understand how anything works. Or they can make software malleable enough that people can construct temporary tools for their own questions. Litt’s optimism rests on the second possibility.
If an agent can build a simulator in minutes, the cost of making a personal instrument for thought falls sharply. A historian can ask for a model that lets her vary an assumption in a population study. A product team can create a live map of how a workflow behaves under different constraints. A developer can visualise a queue, an interpreter or a migration. None of these tools needs to become a commercial product. Their value lies in making an idea manipulable.
That is not the same as saying that generated interfaces are automatically educational. A visualisation can mislead. A simulation can hide the assumptions it claims to reveal. An agent can create a beautiful playground that encodes the wrong model. Human understanding remains the judge. The difference is that the human now has a cheap way to construct an object against which understanding can be tested.
Kay’s old blueprint also explains why the debate about agents cannot be settled by asking whether they make programmers faster. A personal computer was valuable because it enlarged what a person could think and make, not because it reduced keystrokes. The best agent may be the one that helps a user form a sharper question, build an unusual experiment or see a relationship that was previously inaccessible.
In that sense, the return of the programmable computer would not mean returning to a world in which everyone writes traditional code. It would mean returning to a world in which more people can shape the tools they use. Agents may supply the implementation, but the human must still decide what world is worth building and what rules should govern it.
What the Agentic Workplace Must Stop Rewarding
Most organisations say they want innovation, but their operating systems reward visible completion. Tickets closed, features shipped and pull requests merged are easy to count. Understanding is slower and harder to display. It appears in the quality of a question, the accuracy of an estimate, the ability to explain a failure and the speed with which a team can adapt when the original plan proves wrong.
AI intensifies that mismatch. If a manager rewards the number of tasks completed, employees have a reason to accept generated output as quickly as possible. They may avoid the uncomfortable work of reading, experimenting and challenging the agent because that work does not appear in the dashboard. The organisation then congratulates itself on efficiency while transferring the cost into future maintenance.
A better workplace would treat explanation as part of delivery rather than as documentation added after delivery. A significant agent-generated change should arrive with a plain-language account of the system, the design decision, the alternatives and the failure modes. A review should ask whether the team can explain the change, not merely whether an automated check passed.
Managers should also distinguish between different kinds of software work. AI may be highly effective on a well-scoped greenfield task, a repetitive transformation or a throwaway prototype. It may be less effective when the repository contains tacit rules, high standards and years of accumulated context. Productivity policies that treat all tasks as interchangeable will mistake the tool’s strongest setting for a universal law.
The early METR study is instructive here not because its 19 per cent result can be copied into every engineering department, but because it forced a distinction between benchmark success and real-world usefulness. A benchmark may ask whether a model can solve a bounded problem. A team must ask whether the change fits a system, satisfies human expectations and remains understandable after the agent has left.
There is no need for ritual distrust. Humans make bad changes too. The point is to preserve the conditions under which people can notice their own mistakes. A team that cannot explain its software has weakened its internal safety system, even if the current build is passing.
The most sensible performance measure may therefore be a compound one: how much useful work was delivered, how much shared knowledge was created and how easy the result is for another competent person to change. Agents can raise the first number while lowering the other two. A serious organisation has to watch all three.
A Working Discipline for Teams That Use Agents
The practical discipline that follows from Litt’s argument is demanding but not complicated. Every substantial agent-generated change should begin with an explanation of the relevant system, not with a request for more code. The explanation should identify the moving parts, describe the intended behaviour in ordinary language and state what the agent believes it changed. The human should then look for missing assumptions before looking for stylistic imperfections.
Once the change exists, the team should force a distinction between reading and knowing. A short quiz, a prediction exercise or a request to explain the change without opening the diff can expose whether the mental model is real. If the answers are vague, the work is not ready for hand-off. The remedy is not to guess harder. It is to ask for a better explanation or build a small environment that makes the hidden behaviour visible.
When a subsystem is confusing, the default response should not be another page of prose. Ask the agent for a narrow microworld: a toy implementation, a state visualiser, a step-through debugger, a timeline or a simulation with adjustable inputs. The artefact should be designed around the specific question that is blocking progress. A ten-minute experiment can reveal more than another hour spent reading an account that never responds to the reader’s uncertainty.
Teams should also keep the work in a common space. The prompt that generated a major design decision, the explanation of the result and the unresolved questions should be visible to the people who will maintain the system. Private chats are acceptable for rough exploration. They are a poor home for knowledge on which a group will depend.
At regular intervals, each engineer should test the health of the project’s mental model. Could a new colleague be given a clear account of the architecture without being told to read the entire repository? Could the team explain why the most unusual decisions exist? Could someone predict what would happen if a central assumption changed? A failure does not prove incompetence. It is an early warning that cognitive debt is accumulating.
Finally, the time saved by generation should be spent on learning rather than filled with more generation. That may sound inefficient to a culture trained to value constant output. It is the opposite. Reading, questioning and experimenting are what convert a machine’s speed into a person’s capability. If they are omitted, the organisation gets a faster way to produce work it may soon be unable to understand.
The dividing line in agentic engineering will not be between people who use AI and people who refuse it. It will be between teams that use it to deepen their contact with the systems they build and teams that use it to avoid that contact. The first group will treat agents as instruments for making ideas tangible. The second will discover, later and at greater cost, that a system can be automated long before it can be understood.
References
- AI Engineer, “Understanding is the new bottleneck — Geoffrey Litt, Notion”, YouTube, July 2026: https://youtu.be/WkBPX-oDMnA ↩
- METR, “Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity”, 10 July 2025: https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/ ↩
- Simon Willison, “How Generative and Agentic AI Shift Concern from Technical Debt to Cognitive Debt”, 15 February 2026: https://simonwillison.net/2026/Feb/15/cognitive-debt ↩
- Andy Matuschak, “Why books don’t work”: https://andymatuschak.org/books ↩
- Andy Matuschak and Michael Nielsen, “Quantum Country”: https://quantum.country ↩
- “Seymour Papert”, Wikipedia, accessed 28 August 2026: https://en.wikipedia.org/wiki/Seymour_Papert ↩
- Notion, “Notion 3.6: External Agents, HTML blocks, and more”, 1 July 2026: https://www.notion.com/releases/2026-07-01 ↩
- Alan Kay, “A Personal Computer for Children of All Ages”, ACM National Conference, 1972: https://www.mprove.de/visionreality/media/kay72.pdf ↩
This post contains affiliate links. If you purchase through these links, I may earn a commission at no extra cost to you.
Leave a Reply