Proof · real run
One command. 2h 45m hands-off. Approved — receipt included.
The other proofs on this site are recorded or golden demos, and say so. This one is not: a real feature of Orcho's own engine, planned, implemented, reviewed, repaired and accepted in a single run — by Orcho. Below is the receipt the run printed, trimmed for length with the internal identifier and private price figures removed.
2h 45m hands-off, from one command to approved
7 × 5 review × repair rounds, converged to 0 active findings
1 P1 caught at plan validation, fixed on record
~95% of input was cache reads in the recorded run
orcho run · delivery receipt real spend
────────────────────────────────────────────────────────────
[DONE] Pipeline complete
────────────────────────────────────────────────────────────
✓ plan=ok | validate_plan=ok | implement=ok
| review_changes=ok | repair_changes=ok | final_acceptance=ok
• Task: Feature: compressed line-by-line summary grammar
for "--output summary"
• Evidence
• Tasks: 3 planned · 3 completed · 0 failed · 0 incomplete
• Release: approved
• Review findings: 1 (P1=1) | resolved: 1 (P1=1) | active: 0
• Run findings: 4 — attestation: 3 resolved · handoff: 1 resolved
• Open risks: none
• Scope expansion risk: 14 files flagged — unverified · no-explanation
• Verification gates:
↳ pre-final auto-run: 5 ran / 5 pass
↳ blocking (require): broad-non-e2e, verification-unit, cli-sdk-unit
↳ warning (warn): env-provenance, lint — shipping allowed by policy
• Agent advice: calls=1 · applied_retries=1
• Usage: input dominated · ~95% cache-read
Time: 2h 45m | Rounds: 5
• Usage by phase:
↳ plan attempts=2
↳ validate_plan attempts=2
↳ implement attempts=2 (96% cache-read)
↳ review_changes attempts=7 (91% cache-read)
↳ repair_changes attempts=5 (96% cache-read)
↳ final_acceptance accepted
• Sessions: 8 — Claude implements (claude-opus-4-8),
Codex reviews (gpt-5.5); reviewer peak context 69% A feature of Orcho's own CLI, shipped to the public orcho-core repository through Orcho itself. The receipt is trimmed for length; local paths, the internal run identifier and private price figures are removed.
What to look at
- In this profile the reviewer is not the author: a Claude worker implements, a Codex worker reviews. Seven review rounds and five repair rounds converge to zero active findings — the loop is control flow, not ceremony.
- The engine reports on its own agents: 14 files the worker touched but never declared are flagged in the receipt, even though release was approved. Scope honesty survives success.
- Usage is input-dominated and roughly 95% of that input is cache reads. Accounting stays attributed per phase and per subtask, so the operator can see where the workload lives without confusing estimates with a provider invoice.
- Mid-run, the engine's advisor pushed one retry on its own and the stuck attempt resolved without a human touching the run.