Measured, not claimed

Benefits, failures and unknowns keep their denominators.

AIWorkHub separates structural bytes, provider tokens, elapsed time, cost, runtime reliability and manager-accepted quality. Favorable results never hide contrary evidence.

Evidence matrix

SurfacePopulationMeasured resultGrade
Release qualityPython + extension1,944 Python tests passed; 29 extension test files passed; wheel, sdist and VSIX checks passedVerified
Source Graph gate7 current gated tasks7/7 live and satisfied; 13 calls, 0 failures, 846 hits; p50 15.024 msVerified snapshot
Tool-use cohorts170 observed runsReview-ready: missing 7.2%; live single-stage 26.7%; continuous 33.3%Association, not causality
Focused edit shape3 authenticated receipts531 replacement bytes / 31,998 file bytes; 1.66%; 60.26× structural ratio; 0 old bytes re-emittedStructural only
Focused edit A/B2 paired Codex runsHistorical capped A/B observation: 27.5% fewer total tokens, 24.9% fewer output tokens and 21.5% less time; pair 1 used mismatched 20k/200k token ceilingsNot eligible for a causal or product-savings claim; uncapped matched rerun required
Read behavior5 trace-visible tasks4/4 recognized reads bounded; 0 exact/overlap rereads; 15,820 bytesSmall Codex-only pilot
Legacy Context Bundle v1156 tasks on 0.8.81300,419 net bytes added; 20.0% expansion after envelope overheadVerified historical baseline; not current v2 behavior
Context Bundle v2 structural fixOne deterministic same-evidence fixture849 legacy bytes → 600 nested bytes; 29.329% structural reductionShipped in 0.8.82; equivalent live fleet remeasurement pending
Callbacks271 events176 delivered, 95 superseded, 0 dead letters, backlog 0Verified snapshot
Manager quality gate216 decisions25 accepted, 191 rejected; 11.6% acceptanceVerified selectivity
Usage/cost110 records196.35M tokens; $63.82 observed cost; 89 records and 121.14M tokens have unknown costIncomplete cost ledger
Provider parser65 retained real runs63/65 matched the independent terminal aggregate; 2 differences were reference-extractor defectsVerified runtime audit
Routing economics36 costed Claude runsOpus was 19% of tokens and 42.9% of $73.71 cost; same volume at observed Sonnet rate differs by a possible $20.83Unrealized counterfactual, quality parity pending

The measured economic lever is routing, not more cache tuning

ModelRunsCache hitObserved $/M tokensCost
Claude Sonnet 52197.2%$0.658$40.31
Claude Opus 4.8996.4%$1.929$31.61
Claude Haiku 4.5695.0%$0.238$1.79
Codex CLI2974–97% per runUnknownUnknown

Claude caching was already near saturation. Opus cost 2.93× Sonnet and 8.11× Haiku per token in this cohort. The possible $20.83 difference is not realized savings: the next required measurement is observed model × manager acceptance, because a quality advantage may justify the premium.

Retry and reviewer economics

A frozen canonical ledger snapshot contains 114 attempts across 67 tasks. Forty-seven attempts (41.2%) occurred after a task's first recorded attempt and account for 88.31M tokens. Forty-one retry rows have unknown provider cost; the known subset is $20.89. This locates spend but does not claim those tokens were avoidable or that retries caused acceptance.

Historical topic inference identifies 2 reviewer records (1.25M tokens) and 112 worker records (195.78M tokens). New events persist the role explicitly; the ledger keeps inferred legacy rows visibly separate.

The defensible user benefit

Eligible existing-file edits directly avoid full-file code re-emission when that is the baseline. AIWorkHub also makes model-mix economics operational: premium models can be reserved for hard judgment while bounded throughput moves to lower-cost capable routes, and deterministic work moves offline. Task, model, context, isolated workspace, validation evidence, callback and manager decision stay in one repository-scoped loop. A universal Source Graph token multiplier and quality-adjusted system-wide ROI remain unmeasured.

Adjacent execution and context tools

SystemDocumented strengthBoundary relative to AIWorkHub
AIWorkHubMulti-model task DAG, isolated workers, four context layers, Source Graph, hash-bound focused edits, callbacks and evidence-gated manager reviewIntegrated control plane; broader causal performance benchmark is still incomplete
Graphify~40-language tree-sitter graph, communities, query/path/explain and non-code ingestionRicher graph/media product; not documented as a task-DAG/evidence-review control plane
Serena40+ language LSP/IDE symbol retrieval, editing and refactoringDeeper semantic refactoring; not documented as a multi-model orchestration/review system
AiderMature PageRank-selected repo map with a soft token budget and coding benchmarksPair-programming workflow rather than the same repository operations control plane
ClineBroad coding-agent ecosystem; Kanban worktrees/dependencies, checkpoints and many model providersAgent execution/product environment; AIWorkHub is the repository control layer that can coordinate execution routes with canonical context, evidence and economics

These are documented capability boundaries, not a competitor leaderboard. AIWorkHub does not replace coding models or need to claim better graph quality than Graphify, deeper refactoring than Serena, stronger coding scores than Aider or a broader standalone agent ecosystem than Cline. It coordinates such execution and context capabilities as one accountable engineering system.

Claims we refuse to make

  • No universal “× fewer tokens” number for Source Graph.
  • No fleet-wide savings claim from three structural edit receipts.
  • No causal quality claim from observational tool-use cohorts.
  • No system-wide dollar ROI while 121.14M observed tokens lack known cost.

Recompute the evidence

Inspect the system snapshot, provider-routing observation, retry/role observation, paired edit pilot, context-envelope fixture, system checker, routing checker, retry/role checker, context checker and full methodology.