Insights2 min read
Three things a CTO should measure before adopting agentic engineering
Evals, review load and cycle time. If you do not have a baseline for these three before the agents arrive, you will not know whether they helped.
Cliff des Ligneris
Every agentic engineering pilot I have seen fail had the same shape. The team was enthusiastic, the tooling worked, and six weeks later nobody could say whether anything had improved. Not because nothing had changed, but because nobody had measured before they started.
Three numbers fix this. Take them before the first agent touches your repository.
1. Evals
How many of your system’s behaviours are checked by an automated test that a machine can run and a human can read? Not line coverage. Behaviours: “a cancelled booking refunds within the policy window”, “a duplicate invoice is rejected”. Count them.
This is the number that decides how much autonomy you can give an agent. Agents produce code fast and confidently, and some of it is wrong. Evals are the only mechanism that catches the wrong parts without a human reading every diff. If the count is low, the first weeks of a pilot should be spent raising it, not generating features. Teams that skip this step end up with the next problem.
2. Review load
How many pull requests does each senior engineer review per week, and how long does each sit before the first comment? Take the median and the tail.
Agent teams move the bottleneck. Writing code stops being the slow part; judging it becomes the slow part. On Even, review queue depth was the first thing that broke when we scaled the agents up, and the fix was a review agent that filters before a human looks. You cannot see that happening without a baseline.
3. Cycle time
From ticket picked up to change in production. Median and 85th percentile, per team, over the last quarter.
This is the number the board will ask about, and it is the one most easily faked. Faster drafting with slower review nets out to nothing. Only the end-to-end time tells you whether the system improved or just moved the work.
Take all three before you begin. Take them again at four weeks. If evals went up, review load stayed flat, and cycle time fell, you have something. If only the first went up, you have a tooling bill.