Engineering practice · 2026

Running AI coding agents under independent review

I direct AI coding agents. I give them a written bar and a separate reviewer, and I'm trying out a local model for the mechanical work. Most of my code is written this way.

Context
One engineer directing AI coding sessions across many repositories
Period
Process since Aug 2026. Local sub-agent stack since Oct 2026, still a trial.
Built with
Claude Code, llama.cpp, Qwen3.6-35B-A3B on an M5 Max with 128 GB
First review pass
19 pull requests, mostly documentation. 133 must-fix findings raised, 116 confirmed, 17 refuted.

The situation

Since March 2026 I've written most of my code with AI coding agents, mainly Claude Code. Every session starts cold. It remembers nothing from the last one, and it can report success on work it never did.

A written bar

  • A task names one outcome, says how you would know it is done, and lists its blockers in a form a script can read.
  • A fact has one canonical home. Every doc carries a status and the date someone last re-checked it.
  • The task, docs and retrospective standards each have a linter that checks shape. The docs say plainly that no linter can tell whether a claim is true.

Every session is told to end by logging what the process cost it, with evidence, or by saying nothing did. Sessions only log. I decide what changes, and no rule enters a process doc without a dated incident behind it. The log holds more than 340 entries from nine weeks, each with a count of how often it recurs. Its first monthly review is overdue.

A local model for the mechanical work

In October 2026 I started moving the mechanical jobs to an open-weight model on my own machine: Qwen3.6-35B-A3B served by llama.cpp on an M5 Max with 128 GB, with headless Claude Code as the agent loop. The setup is days old, so this part is a trial.

I picked the model with a benchmark of the real work on frozen fixtures: a doc review with nine seeded violations, two lint fixes, and a code audit with one seeded bug. The adoption gate was written before the run, and a model that fails a task loses however fast it is. After the first results I tightened the scoring so a killed agent never counts as a success.

Three models ran. The winner finished all three agentic tasks, run together, in 96.8 seconds. A second model also passed all three but found fewer of the seeded violations. A dense 27B model was killed on one task after 29.8 minutes, still looping.

Serving settings came from a sweep too. On an agent-loop workload, total decode speed stayed near 47 tokens a second from two slots to six, so extra slots only split the same throughput. It runs two. A weekly re-evaluation against new models is scheduled.

The local model gets well-specified jobs, each with its own verifier: a lint, a test, or a cap on diff size. It is not given commit, push, deploy or AWS tools. Those stay with the orchestrating session.

Agent work, then independent review An orchestrating session briefs author agents, some on a local model and some on Claude Sonnet. Each job has its own verifier. The results become pull requests. Independent read-only reviewers read them, two refuters test every must-fix finding, confirmed fixes are applied, and the owner merges. Orchestrating session plans, briefs, checks Author agents local model and Sonnet Per-job verifier lint, test, diff-size cap Pull requests 19, mostly documentation Independent reviewers separate agents, read-only Two refuters per finding survives if neither refutes it Confirmed fixes 116 of 133 Owner merges the gate refused the agent
In this pass the author of a change never reviewed it. Reviewers and refuters were separate read-only agents.

The reviewer is not the author

I don't trust a local model to check its own restructures. Left alone on a large file, it cut 431 lines to 78 and lost 71 facts. A mechanical fact check caught it.

In the first full pass, a mixed pipeline of agents, some local and some on Claude Sonnet, produced 19 pull requests. They were mostly documentation work: docs brought to a written standard, architecture files and restructured entry files, plus some cleanup. One read-only reviewer per repository went through them, each a separate Sonnet agent. Every must-fix finding then went to two refuters with different lenses, and it survived only if neither could refute it. The review used 285 agents and took 34 minutes.

Of 133 must-fix findings raised, 116 were confirmed and 17 were refuted. Twenty-three of the confirmed ones were duplicated text: old wording left beside its corrected version.

When I asked the agent session to merge the 19, the permission gate refused it as self-approval. After a second review pass, I merged them myself.

Checking the checkers

In one run, 5 of 6 verifier agents returned a clean verdict without running the check they were assigned. A message relayed into their context had pulled them off it, and the workflow reported zero errors. A stricter prompt did not fix it on the re-run. I caught it by reading which check each result said it had run.

A local-model agent wrote outside its repository, into a shared checkout and a sibling repository. The tool allow-list covers tool names but not paths. An audit of every checkout after the run caught both edits, and they were reverted. A path-scoped rule is still open.

Open to Staff and Lead backend roles

Remote in the US, or hybrid in Austin, Texas. LinkedIn is the fastest way to reach me.