A semantic system behaving differently from deterministic routing is not evidence that it is better. This project exists to answer the harder question: when the two disagree, which decision is actually more correct? That requires independent ground truth, blinding, explicit adjudication, lossless evidence, uncertainty-aware statistics, and a publication policy that remains valid even when the semantic candidate does not improve the control.
System
Routing Semantic v1 is a completed, P3-verified 120-case comparative study. The oracle was independently adjudicated and frozen before inference; P1 verified the evidence path; P2 ran all 120 cases with the oracle absent and joined only afterward; and P3 published a sanitized result after explicit privacy/integrity verification. In the frozen cohort, the semantic candidate was more accurate on task class and interaction required, while semantic-risk error was approximately unchanged. Routing Semantic v2 is preserving the same evidence-first discipline during qualification: after several fail-closed integration corrections, Benchmark 0.6.3 passed the full live IA-1 through IA-11 qualification with sanitized proof of strict schema transport. Before any real v2 label was produced, a 120-case readiness review exposed a schema-scale representation problem; Benchmark 0.6.4 reduced the equivalent model-facing schema from 225,971 to 3,310 serialized characters by grouping identical decision shapes. Because the representation changed, the earlier pass no longer authorizes the real cohort. Fresh 0.6.4 qualification is required before real routing-semantic-v2 A/B/C. Separately, a fresh 30-pair Agentic-Jev SWE-manager v1 study completed with 15/30 correct in both control and treatment; its request audit found 0/20 successful Jev calls met the frozen context-completeness check. The corrective v2 manager treatment moves authoritative context construction into deterministic host code, has passed live qualification on an already-observed task, and is preregistered for a fresh effectiveness cohort that remains unrun.
Contribution
A completed comparative-study and evidence-engineering pipeline that evolves the question from semantic disagreement to independently measured correctness: frozen cases, blinded oracle creation, staged preflight/full-study execution, seam-specific correctness and calibration, request-level accounting, explicit methodological corrections, and a fail-closed public projection verified before publication.
Architecture & engineering model
Why the methodology changed
flow
The machinery follows from a stricter research question; it is not the point of the study by itself.
Does semantic routing behave differently?
→
When it differs, is it more correct?
→
Need independent ground truth
→
Freeze cases + blinded adjudication
→
Freeze oracle before candidate/control comparison
Independent oracle resolution
flow
Adjudicators see only the frozen case evidence and rubric, never treatment outputs.
frozen authoring view
→
independent A
→
independent B
→
C only for A/B disputes
→
two-of-three majority
→
recorded human resolution or explicit unresolved conflict
Evidence-chain correction after oracle freeze
flow
The v1 oracle is preserved as executed. Private reasoning summaries later showed explanatory evidence existed upstream, but the authoritative pass and original resolution workflow had reduced it to labels.
provider execution + private reasoning summaries
→
labels-only authoritative pass
→
human resolution sees case + rubric + A/B/C labels
A changed decision is only a candidate effect. Whether it is beneficial, harmful, or redundant is measured against independent ground truth.
Blinding protects the oracle
Ground-truth adjudicators cannot see deterministic or semantic outputs, probabilities, comparison results, or hidden construction tags.
Correctness is not operational policy
Raw semantic quality is measured separately from confidence thresholds, fallbacks, lifecycle rules, and the final applied decision.
Unresolved evidence stays unresolved
A genuine three-way conflict can remain explicitly unresolved; the study does not invent a label to preserve sample size.
Methodological shortcomings are evidence too
The authoritative v1 pass omitted structured justifications even though private reasoning summaries survived upstream. The frozen cohort is not rewritten; the corrective response is a versioned future contract with provider-neutral structured justification and end-to-end persistence checks.
Publication is a separate evidence contract
A valid private oracle was not automatically safe to publish. P3 uses a deliberately redacted public projection and an independent verifier for allowlist, integrity, cohort identity, and privacy boundaries.
Freeze executable surfaces, not just version labels
A runtime lock is only reproducible when it binds the result-affecting executable surface and rechecks it before qualification or execution. V2 hardening happened before new treatment outcomes, so it changes the readiness boundary rather than the result.
Publishability is outcome-neutral
The study succeeds when its evidence is defensible and reproducible, even if the semantic candidate is equal to or worse than deterministic control.
The completed study measures deterministic and semantic routing against independently frozen ground truth, then publishes a separately verified public projection. The result is heterogeneous by seam: semantic-candidate accuracy was higher for task class and interaction required, while semantic-risk ordinal error was approximately unchanged.
Frozen cases
120
public-safe shared corpus
Decision seams
3
task class · interaction required · semantic risk
Minimum eligible / seam
100
preregistered publication threshold
Paired bootstrap
10,000
deterministic resamples · 95% confidence
Oracle path
Frozen + validated
qualified DeepSeek adjudication → dispute-only C → recorded human resolution
Oracle API baseline
$0.07
36 DeepSeek-4.1-Flash requests · 726,263 tokens · through P0B freeze
Keep this baseline separate from P1/P2 Jev treatment accounting
Adjudication
P0A/P0B complete
DeepSeek-Flash cohort completed A/B/C plus recorded human resolution
Do not relabel the frozen oracle; preserve the completed cohort as baseline ground truth
Decision provenance
Diagnosis corrected
Authoritative A/B/C pass omitted structured justification; private Inspect reasoning summaries survived upstream but were unavailable to original human resolution
v2 requires provider-neutral structured justification plus round-trip persistence and explicit session-usage scope
Model vs policy
Separated
Raw semantic correctness can be measured independently of thresholds/fallback
Report both layers without conflating them
P1 instrumentation
Passed
Live evidence path verified on 8 development-only cases with 0 exclusions
Retain as the preflight pattern for future cohorts
P2 outcome
Complete
Higher candidate accuracy on task class and interaction required; semantic-risk MAE approximately unchanged
Keep reporting seam-specific; do not collapse into a universal winner
Keep request-level efficiency separate from incomplete provider-cost evidence
Calibration edge case
Corrected + disclosed
One 0.99-mass vector in each multiclass seam; raw evidence preserved and derived calibration copy normalized only
Move probability-mass validation into future preflight/reporting invariants
P3 publication
Verified + published
Private-oracle disclosure path was replaced before release by a redacted public projection with independent verification
Keep private and public evidence contracts separate
Public claims
Seam-specific v1 findings published
No universal Jev superiority, downstream quality causality, cost advantage, or material semantic-risk improvement supported
Require replication/new version for broader claims
Problem
The original question — whether TypeSafe/Jev disagrees with deterministic routing — measures behavioral difference, not correctness.
Successful provider calls, confidence values, and agreement rates cannot establish that a semantic decision is right without an independent reference.
A valid comparison also has to avoid leaking candidate/control outputs into ground-truth creation and must preserve disagreements or unresolved cases rather than quietly forcing a label.
Model quality and host policy are different questions: a semantic answer can be correct even when current thresholds reject it, or incorrect even when policy accepts it.
Technical Highlights
Frozen 120-case public-safe inference corpus with all three routing seams oracle-eligible
Independent oracle authoring view containing only case text, supplied metadata, eligibility, and the frozen decision rubric
Three routing seams: task class (Choice), interaction required (Noul), and semantic risk (Score)
A/B independent adjudication on the exact same hash-pinned blinded artifact
Independent C adjudication only for A/B-disputed cases and seams, without seeing A/B labels
Completed A/B/C oracle adjudication with recorded human resolution for genuine three-way conflicts
Oracle-baseline provider accounting through freeze: $0.07 USD, 36 DeepSeek-4.1-Flash API requests, 726,263 tokens
P1 verified oracle separation, privacy boundaries, probability persistence, unique request IDs, and single-count request-level accounting
Post-freeze evidence audit: the authoritative v1 A/B/C pass retained labels but not structured per-decision justifications; later private-log audit found provider reasoning summaries upstream, but they were not available to the original resolver and were not retrofitted into the frozen oracle
Published seam result: task class 81.67% semantic vs 52.50% control; interaction required 91.67% vs 71.67%; both paired 95% accuracy-difference intervals remained above zero
Semantic-risk result remained approximately unchanged: MAE 0.28975 semantic vs 0.29167 control, with the paired interval spanning zero
Change analysis: interaction required had 30 beneficial changes and 6 harmful changes; task class had 38 beneficial changes, 3 harmful changes, and 10 changed-but-both-wrong cases
Calibration: interaction-required Brier 0.07741 and ECE 0.14942; task-class multiclass Brier 0.27346 and log loss 1.80242; semantic-risk multiclass Brier 0.32484 and log loss 0.53531
Outcome-neutral publication policy: evidence quality is the gate, not a favorable semantic result
Routing Semantic v2 pre-live hardening binds and rechecks result-affecting executable surfaces including the host-side Jev implementation, exact TypeSafe SDK, Inspect/SWE runtime, Codex platform, Docker/Compose and resolved sandbox image identity, and the benchmark Inspect harness
V2 qualification preserves strict whole-completion parsing, schema/provenance identity, and the frozen v1 provenance shape separately from stronger scoped v2 provenance
Benchmark 0.6.3 produced the first full live IA-1 through IA-11 pass with sanitized ingress proof: 15 captured /v1/responses requests carried json_schema + strict=true with schema hashes matching persisted artifacts
Before any real label, real-cohort review measured the 120-case per-case schema at 225,971 serialized characters versus 3,310 for the eligibility-grouped equivalent — a 98.54% representation reduction with all 120 case identities preserved
Benchmark 0.6.4 records the grouped strategy as eligibility-grouped-case-enum/v1 and rejects the older qualification for real runs because the model-facing schema implementation changed
Current routing-semantic-v2 boundary: fresh 0.6.4 IA-1 through IA-11 qualification is required before real A/B; no real v2 A/B/C result is claimed
Agentic-Jev SWE-manager v1 completed 30 fresh paired tasks with 15/30 correct in both arms (0.0 pp observed difference; paired 95% bootstrap interval -16.7 to +20.0 pp); the treatment was Jev availability, not forced use
Jev was called on 20/30 v1 treatment attempts, but 0/20 successful calls met the frozen context-completeness diagnostic; exact provider-dispatch requests were preserved, localizing the loss to pre-dispatch request construction
The versioned manager v2 treatment uses deterministic request construction plus task-specific completeness policy, has passed live qualification on an already-observed task, and remains preregistered/unrun on its fresh effectiveness cohort
System Design
One frozen case is evaluated by deterministic control and one batched TypeSafe/Jev request projected into three decision observations.
The inference runner cannot receive oracle labels. The oracle is a separate artifact and is first joined after inference.
A and B label the same frozen authoring view independently. C receives only disputes. Genuine three-way conflicts require a recorded resolution or remain explicitly unresolved.
The completed oracle-baseline execution used the OpenAI Codex harness through codex-lb to DeepSeek-4.1-Flash; its cumulative provider/API accounting is tracked separately from P1/P2 TypeSafe/Jev traffic.
Raw TypeSafe evidence, Agent-Workflow's policy candidate, and the applied deterministic decision remain separate layers so model quality is not conflated with current operational thresholds or fallback rules.
Correctness and calibration are evaluated per seam; provider overhead is counted once per batched request rather than copied across three decision totals.
Before the full study, P1 exercised the same evidence path on 8 development-only cases and verified request/observation counts, probability persistence, oracle absence, privacy, runtime identities, and accounting invariants.
A future oracle cohort must emit compact structured decision justification in the same independent pass as each label and prove those fields survive persistence/export before real adjudication begins.
P2 reporting preserves raw probability evidence; when one vector in each multiclass seam summed to 0.99, only the derived calibration copy was normalized and the deviation was disclosed.
P3 uses a separate public oracle projection rather than treating the valid private frozen oracle as automatically publishable.
Benchmark tooling executes and sanitizes the experiment; comparative-eval owns comparison meaning; public results are published separately with hashes, identities, limitations, and manifests.
For v2, model-facing schema strategy is part of qualification identity: the 0.6.3 pass remains valid historical evidence, but 0.6.4's grouped real-cohort representation requires a fresh qualification before provider calls are admitted.
Deterministic Authority
Agent-Workflow remains application and lifecycle authority. Comparative mode keeps deterministic control authoritative while recording the semantic candidate as shadow/counterfactual evidence.
Comparative-eval does not schedule work, select models, mutate lifecycle state, or apply candidate decisions.
The oracle is independent evidence, not another source of operational authority.
Result-affecting changes to labels, rubrics, metrics, exclusions, or question meaning require a new study or dataset version.
Where Models Fit
TypeSafe/Jev supplies typed Choice, Noul, and Score evidence for the three registered routing seams.
Probability and confidence are evidence to measure and calibrate, not permission to perform host actions.
The study measures raw semantic correctness against the oracle separately from whether current Agent-Workflow policy would accept, reject, or fall back from that evidence.
Trust, Failure & Recovery
Oracle adjudicators must not see deterministic outputs, TypeSafe/Jev outputs, probability evidence, construction tags, comparison results, or other outside system knowledge excluded by the frozen authoring view.
Every case remains accounted for as eligible, failed, or excluded with a versioned reason; provider failures remain reliability evidence rather than disappearing from denominators.
If A/B disagree, C independently labels only the disputed case/seam. If A/B/C produce no majority, human resolution is recorded or the conflict remains explicitly unresolved.
The authoritative P0B pass preserves labels but not structured per-decision justifications. Human resolution therefore used the frozen case, supplied metadata, frozen rubric, and independent A/B/C labels. A later audit found provider reasoning summaries in private Inspect logs, but those summaries were supplementary execution evidence, were unavailable to the resolver, and are not being retrofitted into the frozen oracle.
The v1 oracle remains immutable as executed. The corrective design belongs in a subsequent protocol/contract version rather than a retroactive rerun presented as historical evidence.
Raw provider HTTP evidence and credentials remain private. P3 publishes a redacted oracle projection that retains final labels plus resolution status/method while removing adjudicator identities, individual A/B/C votes, pass hashes/timestamps, human-resolution rationale/participants, and private reasoning summaries.
The P3 verifier independently checks the exact publication allowlist, SHA manifest integrity, full frozen-sample identity, public-oracle projection contract, privacy flags, and absence of private adjudication/reasoning fields.
V2 qualification is fail-closed and representation-specific. A later implementation change to the real-cohort schema invalidates the earlier passing qualification for authorization purposes even though the earlier pass remains preserved as historical evidence.
The preregistered study specification freezes the corpus target, three decision seams, sample thresholds, exclusion policy, oracle policy, metrics, and deterministic statistical policy before live comparative results are available.
Methodology and study-readiness evidence only. This does not establish that TypeSafe/Jev improves correctness, latency, cost, or any other outcome.
After IA-1 through IA-8 qualification, independent A/B passes labeled the exact blinded authoring view, C received only disputed seams, genuine three-way conflicts went through recorded human resolution, and the resulting oracle was frozen and validated against the 120-case corpus.
This establishes the independent ground-truth baseline later used for scoring. The oracle was frozen before treatment inference and joined only after the full P2 run completed.
Operator-recorded provider/API accounting for initial qualification and adjudication through P0B oracle freeze. Execution path: OpenAI Codex harness -> codex-lb load balancer -> DeepSeek-4.1-Flash API; study tooling recorded the routed alias openai-api/codex-lb/deepseek-flash.
Oracle-baseline construction only. Excludes P1 TypeSafe/Jev smoke, P2 comparative inference, and P3 publication. Raw billing/provider traffic is not committed, so these figures are not used to infer per-token pricing or comparative efficiency.
Post-freeze evidence-chain review first identified that the labels-only authoritative pass omitted structured per-decision justification. Follow-up audit of retained private Inspect logs found contemporaneous provider reasoning-summary events, establishing that explanatory evidence existed upstream but was never promoted into the pass or original human-resolution workflow.
Methodology and evidence-completeness finding only. The private summaries are supplementary execution evidence, not hidden chain-of-thought, were not available to the original resolver, and are not used to rewrite the frozen v1 oracle.
All frozen cases ran with the oracle absent from inference. The frozen oracle was joined only after inference completed. Seam-specific correctness, calibration, disagreement, reliability, latency, and token evidence were then reported under the preregistered policy.
Bounded routing study. Task-class and interaction-required accuracy were higher for the semantic candidate in this cohort; semantic-risk MAE was approximately unchanged and its paired interval spans zero. No universal-superiority, downstream software-quality, or treatment-cost conclusion follows.
Calibration was computed from the persisted semantic probability evidence after the post-inference oracle join. One probability vector in each multiclass seam summed to 0.99; raw evidence was preserved and only the derived calibration copy was normalized.
Calibration evidence for the frozen routing-semantic-v1 cohort. Task-class log loss was 1.80242 and semantic-risk log loss was 0.53531. These metrics do not support a universal model-quality conclusion.
120/120 success · 102,043 tokens · p50 0.1468 s · p95 0.2206 s
Request-level accounting counted the single batched semantic provider request once per case rather than once per decision seam. The run recorded 91,445 input tokens, 10,598 output tokens, and 102,043 provider tokens total; latency was p50 0.1468 s, p90 0.1865 s, and p95 0.2206 s.
Provider request reliability, token, and request-duration evidence for P2 only. Provider cost evidence is incomplete, so these measurements do not support a treatment-cost conclusion.
The publication verifier checked the exact file allowlist, SHA manifest, 120-case/360-observation cohort identity, redacted public oracle projection, privacy flags, and absence of private A/B/C votes, human-resolution rationale/identity, and private reasoning summaries.
Publication-integrity evidence. The public result is a deliberate projection of the valid private evidence bundle rather than a verbatim copy of the private frozen oracle.
A deterministic development-only subset of the frozen corpus exercised the live TypeSafe/Jev path. Verification checked one request per observed case, three seam observations per case, unique request IDs, persisted semantic probabilities/distributions, explicit failures, oracle absence during inference, privacy-safe persistence, single-count request accounting, and frozen routing/v2 plus routing-state/v2 identities.
Instrumentation and evidence-integrity evidence only. The smoke is deliberately below the preregistered study threshold and does not support any TypeSafe/Jev comparative-effectiveness conclusion.
PASS · IA-1 through IA-11 · 15 sanitized ingress requests
An archived Benchmark 0.6.3 retry completed the strengthened qualification. Sanitized loopback capture verified structured-output requests reached Codex-LB ingress with json_schema, strict=true, and schema hashes matching persisted primary/C artifacts; the qualification manifest reported qualified=true and all eleven gates passed.
Qualification evidence for the 0.6.3 schema implementation only. It does not authorize the later 0.6.4 grouped schema and is not a real-cohort oracle result.
Before any real label, deterministic reconstruction compared the per-case anyOf schema with an eligibility-grouped representation. All 120 frozen cases share the same three decision seams, so grouping removes repeated record shapes while retaining the exact allowed case-ID set and unchanged deterministic post-generation validation.
Representation and readiness evidence only. It does not establish that the larger schema would have failed at a provider, and it deliberately invalidates the 0.6.3 qualification for 0.6.4 authorization.