An Eight Mile tool · private

AI ARB

An architecture review board for one code change, on demand. One model plans the change against your repository, a second attacks the plan with the code cited, and they argue it out in rounds until the plan settles. What comes back is one plan, a plain-English record of every objection, and only the decisions that are yours to make.

Runtime

Python 3.11+ · LangGraph

Agents

Claude Code · read-only

Defaults

4 rounds · $60 a run

Tests

1,000+ · no API calls

one run

draft → object → rule

UP TO 4 ROUNDSSCANa fact sheet, no modelPLANNERdrafts the planCRITICraises objectionsSCORERminor ones skip debatePLANNERfixes or rebutsCRITICrules on each answerVERIFIERchecks disputed factsJUDGE · YOUleftovers, by typeEXPLAINERa plain-English ledger

Why argue a plan before it is built

Four rules the debate runs under.

Ask one model for a plan and it is confident. Show it to a second and it finds twenty faults. Relay those back and the first defends itself, and you end up refereeing an argument you may not be able to judge. AI ARB runs that argument for you, under rules.

Every objection cites the code

No complaint without a line number.

The critic must point at the file or the plan line it is objecting to. An objection that cites nothing is refused before it reaches the debate.

“You’re right” is not a resolution

An objection closes on a change or on evidence.

The planner cannot talk its way out. Either the plan changes, or cited evidence shows the problem does not hold. Anything else stays open and goes another round.

Facts are checked, not argued

Disputes about the code go to the code.

When the two sides disagree about what the system does, a verifier reads the repository and settles it. Nobody wins a factual point by arguing it well.

Trade-offs come to you

Priorities are yours, and it knows it.

Fairness against speed, cost against availability: choices like these come back as plain-English questions with each side’s last word, rather than being settled quietly.

A real run

Seven objections, one round, $8.64.

A waiting room for TicketOps, the stadium ticketing product in our case studies, planned by AI ARB on its API repository on 6 October 2026. Quoted from the run’s own record: the brief, the figures, the plan’s milestones before and after, and three of the objections as the explainer wrote them up.

The brief

“Add per-event virtual waiting rooms, built as a new queueing service (queueops) that ticketops integrates with.”

4

Questions first

7

Objections

6 + 1 noted

Settled

1 of 4

Rounds

$8.64

Spent

First draft · 361 lines

M0

decision and skeletons

M1

queueops accounts and rooms

M2

the queueops queue engine

M3

ticketops settings, enable flow and sync (no enforcement yet)

M4

join and enforcement in the api

M5

storefront

M6

hardening

Settled plan · 442 lines

M0

decision and skeletons

M1

queueops accounts and rooms

M2

the queueops queue engine

M3

ticketops settings, sync and enforcement (API only)

M4

join, the human check and the key check

M5

storefront and admin panel

M6

hardening

OBJ-2

major · design

Bots could fill the room before the sale

Places are never given back, and login alone does not stop many verified accounts from taking every place before the sale. The first plan left stronger bot controls for later.

Changed the plan

Resolved. The first release now includes a Turnstile human check that fails closed, a per-address cap of 20 enforced together with capacity, and an optional account-age cutoff. Joining only works through ticketops. Places are still not returned, as the user confirmed. The record says a well-funded farm can still take places, and it lists the warning signs and the response.

The critic’s ruling

The bot gap is closed with login, a Turnstile check that fails closed, a per-address cap enforced atomically in join.lua, and an optional account-age cutoff. The IP source is sound because the api has trustProxy enabled. Capacity semantics follow the user's answer, and the remaining farm risk is recorded.

OBJ-4

major · design

Settings cache and key rotation could lock out buyers

A short-lived cache of room settings on each api replica made enabling or disabling a room slow to take effect. Rotating queueops' signing key without updating ticketops could have silently rejected every admitted buyer.

Changed the plan

Resolved. The cache is removed, so changes apply on the next request. Rotation is now two steps: publish the new key first, then switch signing to it. A join is refused with a 503 if queueops signs with a key ticketops does not know. A cron raises a critical alert on a mismatch. Buyers whose tokens come from a known key keep working.

The critic’s ruling

The cache is removed, so a disable takes effect on the next request. Rotation is two-phase, the join checks signingKid against the pinned keys, and a cron raises an alert on a mismatch.

OBJ-3

minor · trade-off

Early arrival still decides who gets a place

Capacity is first come, first served, so the pre-sale shuffle only orders people who already got a place. People who join after the sale opens go to the back.

Noted, not debated

Noted as a minor trade-off and not debated. The plan now states this consequence in the assumptions, the admin panel and the runbook.

The planner asked four questions before drafting, and the answers are not quoted. Models in this run: claude-opus-5-5 as planner, claude-sonnet-5-5 as critic. Run 2026-10-06-add-per-event-virtual-waiting-rooms-2, AI ARB 1.1.0.

What you get back

One plan, the record behind it, and only your decisions.

A run never forces agreement. What the code can settle, it settles; what only you can decide comes to you as a question with both sides’ last word, and everything ends up in four files beside the repository.

plan.md

The settled plan.

Above it, what to settle before building, the rulings and decisions to apply, and the minor notes that were never debated, each also noted under the section it affects.

ledger.md

Every objection, in plain English.

The problem, what would have happened in production, how it ended, and the full debate behind it.

decisions.md

The trade-offs you decided.

Each one with both sides’ last word, so the reason for a choice outlives the meeting it was made in.

run.json

The whole record, for a machine.

Status, objections, costs and the charter the run used. The command prints JSON and sets an exit code, so a script can tell a settled plan from one that is not ready.

Where it stops

What it does not do, and what is not done.

The debate, the checked ledger, scoring, the decision step, resuming, budgets and the four output files are built and tested. What remains is access, other models, and proof at a larger scale.

Private, and staying that way

access

AI ARB is published to the private package registry of its own repository, not to PyPI, and installing it takes a token that can read that registry.

The run quoted here is not yet checked against a human decision

rough edge

Comparing a run on the ticketing waiting room with the decision reached by hand is the one open item on its real-world checklist.

Claude models only

roadmap

Every agent is a Claude Code session. Harnesses for models outside the Claude family are on the roadmap, as are several planners converging on one plan and replaying past changes to see what the board would have caught.

Redaction is best effort

rough edge

Secrets are redacted from the output files, the log and what the commands print. The checkpoints keep the run’s text as it was, so a resumed run carries on from what the agents actually said; they stay in the run folder, out of git.

Large features take much longer

rough edge

The timings the README publishes are for small features with no debate, at under a dollar each. Most of a run’s cost is the planner’s first draft, which grows with the codebase and the feature.

Behind the project

Built by Eight Mile in London, as part of our AI integration work: the same engineers who build and run systems like it for clients.

Next step

Building something that has to reason before it acts?

AI ARB is ours, and it stays private. The engineers who built it, a LangGraph debate of read-only agents with checked transitions and a budget that counts every call, build AI systems for clients too. Tell us what yours has to get right.

Meet the engineers who would do the work