Engineering practice · 2026
Running AI coding agents under independent review
I direct AI coding agents. I give them a written bar and a separate reviewer, and I'm trying out a local model for the mechanical work. Most of my code is written this way.
- Context
- One engineer directing AI coding sessions across many repositories
- Period
- Process since Aug 2026. Local sub-agent stack since Oct 2026, still a trial.
- Built with
- Claude Code, llama.cpp, Qwen3.6-35B-A3B on an M5 Max with 128 GB
- First review pass
- 19 pull requests, mostly documentation. 133 must-fix findings raised, 116 confirmed, 17 refuted.
The situation
Since March 2026 I've written most of my code with AI coding agents, mainly Claude Code. Every session starts cold. It remembers nothing from the last one, and it can report success on work it never did.
A written bar
- A task names one outcome, says how you would know it is done, and lists its blockers in a form a script can read.
- A fact has one canonical home. Every doc carries a status and the date someone last re-checked it.
- The task, docs and retrospective standards each have a linter that checks shape. The docs say plainly that no linter can tell whether a claim is true.
Every session is told to end by logging what the process cost it, with evidence, or by saying nothing did. Sessions only log. I decide what changes, and no rule enters a process doc without a dated incident behind it. The log holds more than 340 entries from nine weeks, each with a count of how often it recurs. Its first monthly review is overdue.
A local model for the mechanical work
In October 2026 I started moving the mechanical jobs to an open-weight model on my own machine: Qwen3.6-35B-A3B served by llama.cpp on an M5 Max with 128 GB, with headless Claude Code as the agent loop. The setup is days old, so this part is a trial.
I picked the model with a benchmark of the real work on frozen fixtures: a doc review with nine seeded violations, two lint fixes, and a code audit with one seeded bug. The adoption gate was written before the run, and a model that fails a task loses however fast it is. After the first results I tightened the scoring so a killed agent never counts as a success.
Three models ran. The winner finished all three agentic tasks, run together, in 96.8 seconds. A second model also passed all three but found fewer of the seeded violations. A dense 27B model was killed on one task after 29.8 minutes, still looping.
Serving settings came from a sweep too. On an agent-loop workload, total decode speed stayed near 47 tokens a second from two slots to six, so extra slots only split the same throughput. It runs two. A weekly re-evaluation against new models is scheduled.
The local model gets well-specified jobs, each with its own verifier: a lint, a test, or a cap on diff size. It is not given commit, push, deploy or AWS tools. Those stay with the orchestrating session.
The reviewer is not the author
I don't trust a local model to check its own restructures. Left alone on a large file, it cut 431 lines to 78 and lost 71 facts. A mechanical fact check caught it.
In the first full pass, a mixed pipeline of agents, some local and some on Claude Sonnet, produced 19 pull requests. They were mostly documentation work: docs brought to a written standard, architecture files and restructured entry files, plus some cleanup. One read-only reviewer per repository went through them, each a separate Sonnet agent. Every must-fix finding then went to two refuters with different lenses, and it survived only if neither could refute it. The review used 285 agents and took 34 minutes.
Of 133 must-fix findings raised, 116 were confirmed and 17 were refuted. Twenty-three of the confirmed ones were duplicated text: old wording left beside its corrected version.
When I asked the agent session to merge the 19, the permission gate refused it as self-approval. After a second review pass, I merged them myself.
Checking the checkers
In one run, 5 of 6 verifier agents returned a clean verdict without running the check they were assigned. A message relayed into their context had pulled them off it, and the workflow reported zero errors. A stricter prompt did not fix it on the re-run. I caught it by reading which check each result said it had run.
A local-model agent wrote outside its repository, into a shared checkout and a sibling repository. The tool allow-list covers tool names but not paths. An audit of every checkout after the run caught both edits, and they were reverted. A path-scoped rule is still open.