An Eight Mile tool · private
AI ARB
An architecture review board for one code change, on demand. One model plans the change against your repository, a second attacks the plan with the code cited, and they argue it out in rounds until the plan settles. What comes back is one plan, a plain-English record of every objection, and only the decisions that are yours to make.
Runtime
Python 3.11+ · LangGraph
Agents
Claude Code · read-only
Defaults
4 rounds · $60 a run
Tests
1,000+ · no API calls
one run
draft → object → rule
Why argue a plan before it is built
Four rules the debate runs under.
Ask one model for a plan and it is confident. Show it to a second and it finds twenty faults. Relay those back and the first defends itself, and you end up refereeing an argument you may not be able to judge. AI ARB runs that argument for you, under rules.
Every objection cites the code
No complaint without a line number.
The critic must point at the file or the plan line it is objecting to. An objection that cites nothing is refused before it reaches the debate.
“You’re right” is not a resolution
An objection closes on a change or on evidence.
The planner cannot talk its way out. Either the plan changes, or cited evidence shows the problem does not hold. Anything else stays open and goes another round.
Facts are checked, not argued
Disputes about the code go to the code.
When the two sides disagree about what the system does, a verifier reads the repository and settles it. Nobody wins a factual point by arguing it well.
Trade-offs come to you
Priorities are yours, and it knows it.
Fairness against speed, cost against availability: choices like these come back as plain-English questions with each side’s last word, rather than being settled quietly.
A real run
Seven objections, one round, $8.64.
A waiting room for TicketOps, the stadium ticketing product in our case studies, planned by AI ARB on its API repository on 6 October 2026. Quoted from the run’s own record: the brief, the figures, the plan’s milestones before and after, and three of the objections as the explainer wrote them up.
The brief
“Add per-event virtual waiting rooms, built as a new queueing service (queueops) that ticketops integrates with.”
4
Questions first
7
Objections
6 + 1 noted
Settled
1 of 4
Rounds
$8.64
Spent
First draft · 361 lines
M0
decision and skeletons
M1
queueops accounts and rooms
M2
the queueops queue engine
M3
ticketops settings, enable flow and sync (no enforcement yet)
M4
join and enforcement in the api
M5
storefront
M6
hardening
Settled plan · 442 lines
M0
decision and skeletons
M1
queueops accounts and rooms
M2
the queueops queue engine
M3
ticketops settings, sync and enforcement (API only)
M4
join, the human check and the key check
M5
storefront and admin panel
M6
hardening
OBJ-2
major · design
Bots could fill the room before the sale
Places are never given back, and login alone does not stop many verified accounts from taking every place before the sale. The first plan left stronger bot controls for later.
Changed the plan
Resolved. The first release now includes a Turnstile human check that fails closed, a per-address cap of 20 enforced together with capacity, and an optional account-age cutoff. Joining only works through ticketops. Places are still not returned, as the user confirmed. The record says a well-funded farm can still take places, and it lists the warning signs and the response.
The critic’s ruling
The bot gap is closed with login, a Turnstile check that fails closed, a per-address cap enforced atomically in join.lua, and an optional account-age cutoff. The IP source is sound because the api has trustProxy enabled. Capacity semantics follow the user's answer, and the remaining farm risk is recorded.
OBJ-4
major · design
Settings cache and key rotation could lock out buyers
A short-lived cache of room settings on each api replica made enabling or disabling a room slow to take effect. Rotating queueops' signing key without updating ticketops could have silently rejected every admitted buyer.
Changed the plan
Resolved. The cache is removed, so changes apply on the next request. Rotation is now two steps: publish the new key first, then switch signing to it. A join is refused with a 503 if queueops signs with a key ticketops does not know. A cron raises a critical alert on a mismatch. Buyers whose tokens come from a known key keep working.
The critic’s ruling
The cache is removed, so a disable takes effect on the next request. Rotation is two-phase, the join checks signingKid against the pinned keys, and a cron raises an alert on a mismatch.
OBJ-3
minor · trade-off
Early arrival still decides who gets a place
Capacity is first come, first served, so the pre-sale shuffle only orders people who already got a place. People who join after the sale opens go to the back.
Noted, not debated
Noted as a minor trade-off and not debated. The plan now states this consequence in the assumptions, the admin panel and the runbook.
The planner asked four questions before drafting, and the answers are not quoted. Models in this run: claude-opus-5-5 as planner, claude-sonnet-5-5 as critic. Run 2026-10-06-add-per-event-virtual-waiting-rooms-2, AI ARB 1.1.0.
What you get back
One plan, the record behind it, and only your decisions.
A run never forces agreement. What the code can settle, it settles; what only you can decide comes to you as a question with both sides’ last word, and everything ends up in four files beside the repository.
plan.md
The settled plan.
Above it, what to settle before building, the rulings and decisions to apply, and the minor notes that were never debated, each also noted under the section it affects.
ledger.md
Every objection, in plain English.
The problem, what would have happened in production, how it ended, and the full debate behind it.
decisions.md
The trade-offs you decided.
Each one with both sides’ last word, so the reason for a choice outlives the meeting it was made in.
run.json
The whole record, for a machine.
Status, objections, costs and the charter the run used. The command prints JSON and sets an exit code, so a script can tell a settled plan from one that is not ready.
Where it stops
What it does not do, and what is not done.
The debate, the checked ledger, scoring, the decision step, resuming, budgets and the four output files are built and tested. What remains is access, other models, and proof at a larger scale.
Private, and staying that way
AI ARB is published to the private package registry of its own repository, not to PyPI, and installing it takes a token that can read that registry.
The run quoted here is not yet checked against a human decision
Comparing a run on the ticketing waiting room with the decision reached by hand is the one open item on its real-world checklist.
Claude models only
Every agent is a Claude Code session. Harnesses for models outside the Claude family are on the roadmap, as are several planners converging on one plan and replaying past changes to see what the board would have caught.
Redaction is best effort
Secrets are redacted from the output files, the log and what the commands print. The checkpoints keep the run’s text as it was, so a resumed run carries on from what the agents actually said; they stay in the run folder, out of git.
Large features take much longer
The timings the README publishes are for small features with no debate, at under a dollar each. Most of a run’s cost is the planner’s first draft, which grows with the codebase and the feature.
The loop
A LangGraph graph, checkpointed at every step.
Thirteen nodes and five routers. The planner and critic each resume their own Claude Code session from round to round, so a later round covers only the open items rather than rereading the repository.
The graph, as the repository draws it
scan → draft_plan → critique → triage ─┬─▶ respond → rule ─┬─▶ verify ─┐
│ ▲ │ ▲ │ │
▼ │ │ │ │ │
clarify (planner's questions) │ │ │ │
│ └────────────┴───────────┤ another round
│ ▼
└───────────────▶ judge → decide → rescore → explain → finaliserouting.py
Decides every branch. Its return types are the graph’s edges.
finalise
The only node that writes files, and only inside .aiarb/.
resume
Reopened items restart the run from triage, with a fresh set of rounds.
Roles and models
Six roles, and the planner is never its own critic.
Each role has a standing brief of its own, and the charter picks its model. It refuses a charter where the critic is the planner’s model: a model reviewing its own plan shares its blind spots.
| Role | Job | Model in this run |
|---|---|---|
Planner | Asks what it must, drafts the plan, then fixes or rebuts each objection | claude-opus-5-5 |
Critic | Raises objections, each citing the code or the plan, then rules on the answers | claude-sonnet-5-5 |
Scorer | Rates each objection’s impact, before and after the debate. It never decides who is right | claude-sonnet-5-5 |
Verifier | Settles factual disputes by reading the code | claude-sonnet-5-5 |
Judge | Breaks deadlocks left after the final round | claude-opus-5-5 |
Explainer | Rewrites the ledger in plain English | claude-sonnet-5-5 |
Read the source
Five excerpts from closed source.
AI ARB is private, and these are the only lines of it published: the graph's edges, the rule that closes an objection, when the debate stops, the check every agent's tool call passes through, and the models the charter defaults to.
aiarb
src/aiarb/graph
src/aiarb/domain
src/aiarb/adapters
src/aiarb/config
src/aiarb/graph/builder.py · 120–135
graph.add_edge(START, "scan") graph.add_edge("scan", "draft_plan") graph.add_conditional_edges("draft_plan", routing.after_draft) graph.add_edge("clarify", "draft_plan") graph.add_edge("critique", "triage") graph.add_conditional_edges("triage", routing.after_triage) graph.add_edge("respond", "rule") graph.add_conditional_edges("rule", routing.after_rule) graph.add_conditional_edges("verify", routing.after_round) graph.add_conditional_edges("judge", routing.after_judge) graph.add_edge("decide", "rescore") graph.add_edge("rescore", "explain") graph.add_edge("explain", "finalise") graph.add_edge("finalise", END) return graph.compile(checkpointer=checkpointer, name="aiarb") # pyright: ignore[reportUnknownMemberType]
The loop as LangGraph runs it. Fixed edges where there is one way on, and a router wherever the debate can branch: back for another round, to the judge, or to you.
src/aiarb/domain/ledger.py · 298–336
def transition( self, to: Status, *, actor: Actor, round: int, note: str, refs: Iterable[Ref] = (), resolution: Resolution | None = None, ) -> Objection: """Move to another status, if the rules allow it.""" rule = TRANSITIONS.get((self.status, to)) if rule is None: raise TransitionError(f"{self.id}: cannot move from {self.status} to {to}") if actor not in rule.actors: raise TransitionError(f"{self.id}: {actor} cannot move from {self.status} to {to}") if rule.types is not None and self.type not in rule.types: raise TransitionError( f"{self.id}: a {self.type} objection cannot move from {self.status} to {to}" ) if to is Status.RESOLVED and resolution is None: raise TransitionError( f"{self.id}: an objection is resolved only by a plan change or cited evidence" ) if to is not Status.RESOLVED and resolution is not None: raise TransitionError(f"{self.id}: only a move to {Status.RESOLVED} has a resolution") if to is Status.NOTED and not self.severity.skips_debate: raise TransitionError( f"{self.id}: a {self.severity} objection must be debated, not noted" ) return self._record( round=round, actor=actor, status=to, severity=self.severity, note=note, refs=tuple(refs), resolution=resolution, )
The closing rule. Every move is checked against a table of who may make it, and a move to resolved without a plan change or cited evidence is refused outright.
src/aiarb/graph/routing.py · 50–92
def after_round(state: RunState, runtime: Runtime[RunContext]) -> AfterRound: """Go another round, hand the leftovers to the judge, or finish the debate. Trade-offs can be waiting for the user once the debate ends without the judge: after the user reopened other items instead of deciding them. """ if not _any_open(state): return "decide" if decisions_needed(state) else "rescore" if judge_needed(state, round_limit(state, runtime.context.charter)): return "judge" return "respond" def after_judge(state: RunState) -> AfterJudge: """Ask the user about trade-offs left for them, or move on to rescoring.""" return "decide" if decisions_needed(state) else "rescore" def judge_needed(state: RunState, max_rounds: int) -> bool: """The debate has items left but must stop: out of rounds, or going in circles.""" return round_limit_reached(state, max_rounds) or stalled(state) def round_limit(state: RunState, charter: Charter) -> int: """The run's round limit: the one it was given, or the charter's.""" return state["max_rounds"] or charter.debate.max_rounds def round_limit_reached(state: RunState, max_rounds: int) -> bool: """The current debate has had its rounds. A debate the user reopened gets a full set of its own.""" return state["round"] - state["debate_start"] >= max_rounds def stalled(state: RunState) -> bool: """No objection has left the debate for `STALL_ROUNDS` rounds. Leaving means being resolved, deferred, sent for a decision, or anything else that ends the argument over it. A rejected answer does not count, even if the plan changed, since the critic was not persuaded. The user reopening items restarts the clock. """ return state["round"] - max(last_progress(state), state["debate_start"]) >= STALL_ROUNDS
When the debate stops: nothing left open, the round limit reached, or two rounds in a row where no objection left the debate. A rejected answer does not count as progress.
src/aiarb/adapters/claude_code.py · 334–349
def refusal(root: Path, tool: str, tool_input: dict[str, Any]) -> str | None: """Why this tool call may not run, or None if it may.""" if tool == ANSWER_TOOL: return None if tool not in READ_ONLY_TOOLS: return f"{tool} is not available: this review may only read and search the repository." for key in ("file_path", "path"): value = tool_input.get(key) if value is None: continue if not isinstance(value, str) or not _inside(root, value): return f"{value!r} is outside the repository; only files inside {root} can be read." pattern = tool_input.get("pattern") if tool == "Glob" and isinstance(pattern, str) and _escapes(pattern): return f"the pattern {pattern!r} reaches outside the repository." return None
Every tool call an agent makes passes through this first. Only Read, Grep and Glob run, and only on paths inside the repository under review.
src/aiarb/config/charter.py · 39–58
class Models(_Section): """The model each role runs on. Any model ID Claude Code accepts.""" planner: ModelId = "claude-opus-5-5" critic: ModelId = "claude-sonnet-5-5" judge: ModelId = "claude-opus-5-5" verifier: ModelId = "claude-sonnet-5-5" explainer: ModelId = "claude-sonnet-5-5" scorer: ModelId = "claude-sonnet-5-5" """Used by the language-model scorer, whether chosen or as Jev's fallback, and to confirm blocker calls.""" @model_validator(mode="after") def _critic_differs_from_planner(self) -> "Models": if self.critic == self.planner: raise ValueError( f"the critic must be a different model from the planner (both are " f"{self.planner!r}): a model reviewing its own plan shares its blind spots" ) return self
The default model for each role, every one a setting in the charter. The one rule enforced on them: the critic is never the planner's model.
Rounds, budget, resuming
Bounded in rounds and in dollars, and never lost.
A run stops when the plan settles, when it runs out of rounds or money, or when it needs you. Each of those leaves it resumable.
Rounds
Up to 4 rounds, then leftovers by type.
A fact the code cannot settle becomes a task before you build. A trade-off comes to you. A circular argument goes to the judge and is marked “ruled, not agreed”. Two rounds in a row with no objection leaving the debate count as a stall, and the leftovers go to the judge early.
Budget
$10 a call, $60 a run, every call counted.
Both caps are charter settings. The run total counts retried and failed calls too. Once it is spent no further call starts, and the run stops with exit code 5, keeps its progress and carries on once the budget is raised.
Resuming
Every step is a checkpoint.
LangGraph checkpoints each finished step to SQLite in the run’s own folder, so a run that pauses for your decisions, fails a call or is stopped with Ctrl-C carries on from its last finished step with aiarb resume.
Reopening
New evidence restarts the debate.
Load-test results after the run? aiarb resume --evidence reopens only the items they bear on, and those are debated again with a full set of rounds of their own.
What it leaves in a repo
One folder, and two files in it are committed.
Agents can only read, and a run writes nothing outside .aiarb/. As a check, each run compares the repository’s files before and after, and warns if anything outside the folder changed while it ran.
.aiarb
.aiarb
.aiarb/runs/<date>-<name>
.aiarb/cache/facts
.aiarb/charter.yaml
committed
How the board behaves in this repository: a model per role, the round limit, scoring, what matters here and how much, rules every plan must keep to, and the budget. Every setting is optional, and an unknown key is refused so a typo fails loudly. This is the README’s example.
models: planner: <model-id> # your strongest model critic: <model-id> # must differ from the planner judge: <model-id> # takes no part in the debate verifier: <model-id> # a smaller model is fine explainer: <model-id> # a smaller model is fine scorer: <model-id> # scores when Jev is not used, and confirms blockers debate: max_rounds: 4 scoring: provider: llm # or: jev, with TYPESAFE_API_KEY set confidence_threshold: 0.7 # below this, a language model rescores Jev's answer priorities: # what matters in this codebase, and how much data_integrity: 5 availability_at_peak: 5 delivery_time: 3 running_cost: 2 constraints: - no new managed services - zero-downtime deploys only budget: per_call: 10 # spend cap per agent call per_run: 60 # the run stops cleanly and stays resumable
.aiarb/.gitignore
committed
Written by aiarb init, so runs and the scan cache never reach the repository’s history.
# aiarb's runs and scan cache stay on this machine. The charter is committed. runs/ cache/
.aiarb/runs/<date>-<name>/plan.md
out of git
The settled plan, with what to settle before building, the rulings, your decisions and the minor notes at the top, each also noted under the section it affects.
.aiarb/runs/<date>-<name>/ledger.md
out of git
Every objection in plain English: the problem, what would have happened in production, how it ended, and the full debate.
.aiarb/runs/<date>-<name>/decisions.md
out of git
The trade-offs you decided, with each side’s last word.
.aiarb/runs/<date>-<name>/run.json
out of git
The full machine-readable record. The run on this page was quoted from one of these.
.aiarb/runs/<date>-<name>/log.jsonl
out of git
One line for each step and each outside call, appended as the run goes: when, how long, on what model, what it spent, and how it failed if it did.
.aiarb/runs/<date>-<name>/checkpoints.sqlite
out of git
LangGraph’s checkpoint of every finished step, with the run’s ID as its thread, which is what aiarb resume carries on from. It keeps the run’s text unredacted, which is one more reason it stays out of git.
.aiarb/cache/facts/<commit>-v<n>.md
out of git
The scan’s fact sheet of the repository, built without a model and cached by commit, so a second run on the same commit skips the scan.
Where it stops
What it does not do, and what is not done.
The debate, the checked ledger, scoring, the decision step, resuming, budgets and the four output files are built and tested. What remains is access, other models, and proof at a larger scale.
Private, and staying that way
AI ARB is published to the private package registry of its own repository, not to PyPI, and installing it takes a token that can read that registry.
The run quoted here is not yet checked against a human decision
Comparing a run on the ticketing waiting room with the decision reached by hand is the one open item on its real-world checklist.
Claude models only
Every agent is a Claude Code session. Harnesses for models outside the Claude family are on the roadmap, as are several planners converging on one plan and replaying past changes to see what the board would have caught.
Redaction is best effort
Secrets are redacted from the output files, the log and what the commands print. The checkpoints keep the run’s text as it was, so a resumed run carries on from what the agents actually said; they stay in the run folder, out of git.
Large features take much longer
The timings the README publishes are for small features with no debate, at under a dollar each. Most of a run’s cost is the planner’s first draft, which grows with the codebase and the feature.
Behind the project
Built by Eight Mile in London, as part of our AI integration work: the same engineers who build and run systems like it for clients.
Next step
Building something that has to reason before it acts?
AI ARB is ours, and it stays private. The engineers who built it, a LangGraph debate of read-only agents with checked transitions and a budget that counts every call, build AI systems for clients too. Tell us what yours has to get right.