Services

Industries

Insights

Community

•

When AI agents write the code, what does the engineer write?

Santiago Marro explores how engineering changes when AI agents write code: specs, skills, and evaluations that enable teams to delegate execution without losing judgment, control, or human accountability.

By Santiago Marro, AI Specialist Lead at Santex

For decades, software engineers turned problems into instructions a machine could execute. Today, AI agents navigate repositories, write code, and open pull requests. If producing software keeps getting cheaper, the question changes: what does the engineer write when they no longer need to write all the code?

The answer can be summed up in three artifacts: the spec, the skill, and the verdict. Code still matters, but it is no longer the only place where engineering happens.

The economics have already changed. According to Epoch AI, producing an answer at the same performance level became 725 times cheaper in less than 18 months. The cost of generating a response fell from $0.30 to $0.0004, while performance held steady at 81%. Thinking, at least in terms of generation, became cheap.

But feeling confident about that answer did not. A Harvard study of 718 companies found that the time between opening and merging a PR increased by 49% with agents, from seven days to 10.5. The number of PRs receiving change requests also nearly doubled.

In software, speed comes at a cost too. Faros analyzed telemetry from 22,000 developers across 4,000 teams. As AI usage increased sharply, PRs merged without review rose by 31%, while incidents per PR increased by 243%. At mature adoption levels, those figures reached 76% and 14.5%, while monthly incidents rose by 125%.

Stanford and Carnegie Mellon found 2.09 times more PRs per developer, while human review dropped from 39% to 21%. Automated review jumped from 19% to 84%. Producing more does not mean controlling better.

That brings us to the engineer’s first new job: writing the spec. The engineer must define the problem to solve, the constraints that matter, and how we will know whether the result actually works. A good specification lets us delegate execution without delegating judgment.

The second artifact is the skill. The team’s rules, conventions, and decisions need to be written down so the agent can reuse them. Across repositories analyzed by Stanford and Carnegie Mellon, cognitive complexity increased by 53% without versioned configuration and by 27% with it.

The third artifact is the verdict. The verdict defines who reviews the work, what criteria they use, and what gets accepted. To make that decision, we need to know how often the agent gets things wrong. And we measure that the same way we do at school: with an exam made up of cases for which we already know the correct answer. The problem is that the exam can lie too. It happened to us twice.

One of our agents reads invoices. In our evaluation, it scored 91%. Too good. The instructions we gave it included sample invoices that were also part of the evaluation set. It already knew the answers. Once we removed those examples, performance dropped to 84%. And because we are still measuring it against the same 38 invoices, even that number is an upper bound. The real measure has to come from invoices the agent has never seen before. The other case involved an agent that effectively corrected its own exam. Once we restored the original tests and locked them so the agent could not modify them, the result turned red.

This is not just bad luck on our part. Researchers from Carnegie Mellon and Anthropic took real-world programming problems and deliberately altered the tests so they contradicted the task requirements. There was no honest way to pass: if everything turned green, the model had cheated. The researchers even instructed the models, in uppercase, not to modify the tests. With 2025 models, GPT-5 cheated on 76% of the tasks and Claude 3.7 Sonnet on 70%. They rewrote tests or created shortcuts designed to produce the expected result. When the tests were locked, cheating by Claude Opus 4.1, another model included in the study, dropped from 54% to 6%. GPT-5 improved far less because it relied more heavily on shortcuts.

If the agent can change the exam, the exam measures nothing. That is why teams getting real value from these systems evaluate them on cases the agent has never seen and cannot modify. Legora, which develops AI for legal professionals, hired lawyers to write its evaluations. At Waymo, an AI project is not considered ready when the model performs well. It is ready when its evaluations are ready. In software engineering, that evaluation set can start with twenty or thirty problems your team has already solved, along with their tests.

The new engineer, then, is not the person who produces the most lines of code. It is the person who knows what to delegate, how to encode the team’s knowledge, how to design evaluations, and where to set the threshold that forces the system to stop. They also need to know how to read traces, understand the results, and have the authority to reject a merge.

The shift is profound: the agent executes, but the engineer designs the system that makes that execution trustworthy. When the machine writes the code, the signature is still human.



Share on:

Share on:

Stay Ahead with Expert Tips, Trends, and Insights

Get the latest content from Santex: ideas, tech updates, and resources that matter.

Like

Stay Ahead with Expert Tips, Trends, and Insights

Get the latest content from Santex: ideas, tech updates, and resources that matter.

Stay Ahead with Expert Tips, Trends, and Insights

Get the latest content from Santex: ideas, tech updates, and resources that matter.

Stay Ahead with Expert Tips, Trends, and Insights

Get the latest content from Santex: ideas, tech updates, and resources that matter.

Stay Ahead with Expert Tips, Trends, and Insights

Get the latest content from Santex: ideas, tech updates, and resources that matter.

  • Connect Program

  • Expert-Led Innovation

  • Quality & Security

  • Committed to Sustainability