Evidence matrix
| Surface | Population | Measured result | Grade |
|---|---|---|---|
| Release quality | Python + extension | 1,944 Python tests passed; 29 extension test files passed; wheel, sdist and VSIX checks passed | Verified |
| Source Graph gate | 7 current gated tasks | 7/7 live and satisfied; 13 calls, 0 failures, 846 hits; p50 15.024 ms | Verified snapshot |
| Tool-use cohorts | 170 observed runs | Review-ready: missing 7.2%; live single-stage 26.7%; continuous 33.3% | Association, not causality |
| Focused edit shape | 3 authenticated receipts | 531 replacement bytes / 31,998 file bytes; 1.66%; 60.26× structural ratio; 0 old bytes re-emitted | Structural only |
| Focused edit A/B | 2 paired Codex runs | Historical capped A/B observation: 27.5% fewer total tokens, 24.9% fewer output tokens and 21.5% less time; pair 1 used mismatched 20k/200k token ceilings | Not eligible for a causal or product-savings claim; uncapped matched rerun required |
| Read behavior | 5 trace-visible tasks | 4/4 recognized reads bounded; 0 exact/overlap rereads; 15,820 bytes | Small Codex-only pilot |
| Legacy Context Bundle v1 | 156 tasks on 0.8.81 | 300,419 net bytes added; 20.0% expansion after envelope overhead | Verified historical baseline; not current v2 behavior |
| Context Bundle v2 structural fix | One deterministic same-evidence fixture | 849 legacy bytes → 600 nested bytes; 29.329% structural reduction | Shipped in 0.8.82; equivalent live fleet remeasurement pending |
| Callbacks | 271 events | 176 delivered, 95 superseded, 0 dead letters, backlog 0 | Verified snapshot |
| Manager quality gate | 216 decisions | 25 accepted, 191 rejected; 11.6% acceptance | Verified selectivity |
| Usage/cost | 110 records | 196.35M tokens; $63.82 observed cost; 89 records and 121.14M tokens have unknown cost | Incomplete cost ledger |
| Provider parser | 65 retained real runs | 63/65 matched the independent terminal aggregate; 2 differences were reference-extractor defects | Verified runtime audit |
| Routing economics | 36 costed Claude runs | Opus was 19% of tokens and 42.9% of $73.71 cost; same volume at observed Sonnet rate differs by a possible $20.83 | Unrealized counterfactual, quality parity pending |
The measured economic lever is routing, not more cache tuning
| Model | Runs | Cache hit | Observed $/M tokens | Cost |
|---|---|---|---|---|
| Claude Sonnet 5 | 21 | 97.2% | $0.658 | $40.31 |
| Claude Opus 4.8 | 9 | 96.4% | $1.929 | $31.61 |
| Claude Haiku 4.5 | 6 | 95.0% | $0.238 | $1.79 |
| Codex CLI | 29 | 74–97% per run | Unknown | Unknown |
Claude caching was already near saturation. Opus cost 2.93× Sonnet and 8.11× Haiku per token in this cohort. The possible $20.83 difference is not realized savings: the next required measurement is observed model × manager acceptance, because a quality advantage may justify the premium.
Retry and reviewer economics
A frozen canonical ledger snapshot contains 114 attempts across 67 tasks. Forty-seven attempts (41.2%) occurred after a task's first recorded attempt and account for 88.31M tokens. Forty-one retry rows have unknown provider cost; the known subset is $20.89. This locates spend but does not claim those tokens were avoidable or that retries caused acceptance.
Historical topic inference identifies 2 reviewer records (1.25M tokens) and 112 worker records (195.78M tokens). New events persist the role explicitly; the ledger keeps inferred legacy rows visibly separate.
The defensible user benefit
Eligible existing-file edits directly avoid full-file code re-emission when that is the baseline. AIWorkHub also makes model-mix economics operational: premium models can be reserved for hard judgment while bounded throughput moves to lower-cost capable routes, and deterministic work moves offline. Task, model, context, isolated workspace, validation evidence, callback and manager decision stay in one repository-scoped loop. A universal Source Graph token multiplier and quality-adjusted system-wide ROI remain unmeasured.
Adjacent execution and context tools
| System | Documented strength | Boundary relative to AIWorkHub |
|---|---|---|
| AIWorkHub | Multi-model task DAG, isolated workers, four context layers, Source Graph, hash-bound focused edits, callbacks and evidence-gated manager review | Integrated control plane; broader causal performance benchmark is still incomplete |
| Graphify | ~40-language tree-sitter graph, communities, query/path/explain and non-code ingestion | Richer graph/media product; not documented as a task-DAG/evidence-review control plane |
| Serena | 40+ language LSP/IDE symbol retrieval, editing and refactoring | Deeper semantic refactoring; not documented as a multi-model orchestration/review system |
| Aider | Mature PageRank-selected repo map with a soft token budget and coding benchmarks | Pair-programming workflow rather than the same repository operations control plane |
| Cline | Broad coding-agent ecosystem; Kanban worktrees/dependencies, checkpoints and many model providers | Agent execution/product environment; AIWorkHub is the repository control layer that can coordinate execution routes with canonical context, evidence and economics |
These are documented capability boundaries, not a competitor leaderboard. AIWorkHub does not replace coding models or need to claim better graph quality than Graphify, deeper refactoring than Serena, stronger coding scores than Aider or a broader standalone agent ecosystem than Cline. It coordinates such execution and context capabilities as one accountable engineering system.
Claims we refuse to make
- No universal “× fewer tokens” number for Source Graph.
- No fleet-wide savings claim from three structural edit receipts.
- No causal quality claim from observational tool-use cohorts.
- No system-wide dollar ROI while 121.14M observed tokens lack known cost.
Recompute the evidence
Inspect the system snapshot, provider-routing observation, retry/role observation, paired edit pilot, context-envelope fixture, system checker, routing checker, retry/role checker, context checker and full methodology.