Omerta
roadmap · 28 Sep 2026

A holdout set for your coding agents. Checks decide, not the model.

Omerta runs Claude Code and Codex as sandboxed workers and accepts a change only when checks the workers can't see say it works.

status private, in development
built for one developer's family projects
current phase B, breadth · front end started early

Where it stands

The same five real open-source bugs, three trials each, identical clones and identical grading for every arm.

AgentSolved
Omerta14 of 15
Bare Codex8 of 15
Bare Claude Code4 of 15

Directional, not a benchmark. Five tasks is a small sample, and Omerta wrote the hidden checks that graded every arm. Phase B fixes that with tasks graded by upstream tests Omerta never wrote.

Slower and pricier per task. An easy fix takes about 70 seconds against about 50 bare. Hard fixes take 10 to 50 minutes, because every change is proven and reviewed.

Workers provably can't read the hidden checks. This was verified live. Every check must fail on the original code and pass on a known-good fix before it counts.

Six phases, each with a gate

Omerta moves on only when the scoreboard shows the gate is met. Each phase says what it is, what's next, and what Omerta can take on once it's done.

A
done

Foundations: speed, isolation, routing

Omerta learned to write calibrated hidden checks, keep workers away from them, review changes by running both versions, route work to the right model and escalate after failures, and do it fast enough to use.

gate easy ≤ 90 s, hard ≤ 25 min, 3 trials, isolation proven → met (70 s easy · 14/15 · isolation verified)

Now it can fix real bugs in C, C++ and Python projects more reliably than the same models on their own.

B
now

Breadth: any common language, any kind of task

Turn a strong bug-fixer on a handful of tasks into a general tool that works on real projects with real dependencies, from a plain description, and prove it on a task set Omerta didn't grade itself.

Done

  • Dependencies installed once per lockfile and run offline in the sandbox (npm, pnpm, Python, Go, Rust); proven on a Next.js app
  • Checks can start the app and drive a real browser

Next

  • 25+ tasks mined from already-fixed upstream issues, graded by the upstream fix's own tests
  • Onboarding from a description alone: Omerta infers what may change and how to test it
  • Features, refactors and dependency upgrades, not just fixes
  • The debugger as a tool the worker can pick up

gate ≥ 90% solved, ≥ 20 points above both bare agents, ≤ 1.5× bare Claude's time → open

After B it can be a daily driver: point it at any of the family's repos with a sentence, and get back a change you can trust, with the evidence.

C
started early

Front end: built to a pinned design

Design quality decides front-end work, so Omerta gets a design authority. You approve a concept sheet, it becomes the pinned contract, and user journeys run in a real browser at phone, tablet and desktop widths. A reviewer compares screenshots against the pinned frames.

Done

  • The app runs inside the sandbox with its own database and demo data; headless Chromium drives it
  • Offline fonts and icons, so screenshots match production
  • Design pins, a visual reviewer, and a front-end route on the strongest model
  • First concept approved: a family caregiving app, "the grove"

Next

  • Finish the first task (the new foundation and the Today screen), judged against bare Claude Code
  • Two more: the five-step check-in, then landing and sign-in
  • A style-locked illustration set for the nature feel
  • Checks against sameness, and conversion checks

gate 3 front-end tasks, every journey passes, rated above bare Claude Code → open

After C it can take a concept sheet and build the product to that taste, screen by screen, with journeys that prove it works.

D
later

Long horizon: days, not minutes

Big goals become milestones and blocks, each with its own checks. Project memory keeps decisions and lessons, the regression suite only grows, and an unattended supervisor watches budgets, notices when it's stuck, and texts a human only when needed.

gate a 48-hour multi-milestone build, killed on purpose, recovers and passes its full suite

After D it can own a weekend project: "bring the mobile app to parity with the web" runs while you're away.

E
later

Parallelism: many workers, no collisions

Shared interfaces are agreed and frozen first. Each worker owns its files under a lease, a merge queue re-proves every merge, and a scheduler keeps the critical path moving.

gate ≥ 1.3× faster at 2 workers with no loss of quality, then 4, then 8

After E it can build a whole app in about a day, or several family projects at once.

F
later · track starts now

Self-improvement: better every week, from its own evidence

Omerta tunes its own prompts, model routes, gates and budgets with A/B tests on tasks it has never seen, and proposes its own work. It never touches what grades it.

gate measured gains on a sealed task split, zero grader regressions, a human approves every prompt or gate change

After F it can improve without being prompted, and ask only for decisions.

Autonomy

Can Omerta ask itself questions, choose its own work and improve itself? Yes, in a bounded form.

Autonomy goes to the side that proposes work. The side that accepts it stays frozen and out of reach.

What the research shows

20 → 50%

A self-rewriting coding agent on a SWE-bench subset, scored by tests it couldn't edit.

17 → 53%

A second self-improving coding agent, under the same condition.

+10 pts

Self-play that invents its own bugs to fix, with every bug proven by executable tests.

faked logs

What self-improvers did when they could reach their grader. One reported over 1000% accuracy.

76%

A frontier model cheating on deliberately impossible coding tasks.

0%

Frontier AI research that labs report running fully on its own today.

How Omerta does it

  1. Reflect. A scheduled job reads the ledger and writes open questions from losses to bare agents, flaky passes, reviewer disagreements, slow or costly runs, and tasks a worker called impossible.
  2. Hypothesize. Each question gets a claim that can be proven wrong and a pre-registered experiment.
  3. Test on sealed tasks. Experiments run on a split the loop never trains on, graded by checks it never wrote.
  4. Promote by tier. Numbers may change themselves; prompts and gates need a human; the grader never changes.
  5. Audit the grader. Planted shortcuts and impossible tasks catch any check that can be gamed. Hard budgets and an automatic kill switch sit above everything.
NeverThe grader, hidden checks, held-out tasks, sandbox, budgets, kill switch
AutomaticNumeric settings such as timeouts, retries and escalation thresholds, in Phase F
Human approvalPrompts, gates, model routes
Frozen for nowThe check writer and calibration, until phases B and D hold

The end goal

For the family: a team you can hand work to

Anyone in the family says what they want in plain words, like "add refill reminders to the care app". Omerta turns that into checks you can read, builds it, proves it, has it reviewed, works overnight when it needs to, and hands back a change ready to merge. It texts you only when a decision is yours.

north star = share of weekly asks merged as delivered, with zero regressions shipped

If it goes open source

The part worth sharing is the acceptance layer: calibrated hidden checks, worker isolation, a reviewer that runs the code, and a scoreboard graded by tests Omerta never wrote. Running agents side by side is already crowded; trustworthy acceptance isn't.

  • First pass the Phase B gate with outside grading
  • Choose the license on purpose
  • Keep family projects and personal specs private
  • Support API-key sign-in, never borrowed consumer logins