Skip to content
Nate G. / Vancouver, WA · US
all work
Published StudyEvaluation EngineeringLLM EvaluationExperimental DesignCalibration

Agent-Workflow Comparative Evaluation

From measuring semantic disagreement to measuring decision correctness against an independently frozen oracle.

Case summary

Purpose

A semantic system behaving differently from deterministic routing is not evidence that it is better. This project exists to answer the harder question: when the two disagree, which decision is actually more correct? That requires independent ground truth, blinding, explicit adjudication, lossless evidence, uncertainty-aware statistics, and a publication policy that remains valid even when the semantic candidate does not improve the control.

System

Routing Semantic v1 is a completed, P3-verified 120-case comparative study. The oracle was independently adjudicated and frozen before inference; P1 verified the evidence path; P2 ran all 120 cases with the oracle absent and joined only afterward; and P3 published a sanitized result after explicit privacy/integrity verification. In the frozen cohort, the semantic candidate was more accurate on task class and interaction required, while semantic-risk error was approximately unchanged. Routing Semantic v2 is preserving the same evidence-first discipline during qualification: after several fail-closed integration corrections, Benchmark 0.6.3 passed the full live IA-1 through IA-11 qualification with sanitized proof of strict schema transport. Before any real v2 label was produced, a 120-case readiness review exposed a schema-scale representation problem; Benchmark 0.6.4 reduced the equivalent model-facing schema from 225,971 to 3,310 serialized characters by grouping identical decision shapes. Because the representation changed, the earlier pass no longer authorizes the real cohort. Fresh 0.6.4 qualification is required before real routing-semantic-v2 A/B/C. Separately, a fresh 30-pair Agentic-Jev SWE-manager v1 study completed with 15/30 correct in both control and treatment; its request audit found 0/20 successful Jev calls met the frozen context-completeness check. The corrective v2 manager treatment moves authoritative context construction into deterministic host code, has passed live qualification on an already-observed task, and is preregistered for a fresh effectiveness cohort that remains unrun.

Contribution

A completed comparative-study and evidence-engineering pipeline that evolves the question from semantic disagreement to independently measured correctness: frozen cases, blinded oracle creation, staged preflight/full-study execution, seam-specific correctness and calibration, request-level accounting, explicit methodological corrections, and a fail-closed public projection verified before publication.

Architecture & engineering model

Why the methodology changed

flow

The machinery follows from a stricter research question; it is not the point of the study by itself.

Does semantic routing behave differently?
When it differs, is it more correct?
Need independent ground truth
Freeze cases + blinded adjudication
Freeze oracle before candidate/control comparison

Independent oracle resolution

flow

Adjudicators see only the frozen case evidence and rubric, never treatment outputs.

frozen authoring view
independent A
independent B
C only for A/B disputes
two-of-three majority
recorded human resolution or explicit unresolved conflict

Evidence-chain correction after oracle freeze

flow

The v1 oracle is preserved as executed. Private reasoning summaries later showed explanatory evidence existed upstream, but the authoritative pass and original resolution workflow had reduced it to labels.

provider execution + private reasoning summaries
labels-only authoritative pass
human resolution sees case + rubric + A/B/C labels
promotion gap documented
future contract: provider-neutral structured justification + persistence preflight

Correctness and policy remain separate

layers

This prevents an application threshold or fallback rule from being mistaken for model quality.

Raw semantic evidence

  • Choice / Noul / Score
  • probability and confidence
  • evaluated directly against oracle

Operational policy

  • confidence thresholds
  • fallback and non-automation rules
  • policy candidate remains inspectable

Applied decision

  • deterministic control remains authoritative
  • application owns side effects and lifecycle
  • comparison does not transfer authority

Publishable study path

flow

Truth is first joined after inference; correctness and request-level overhead are then reported under frozen sample and publication rules.

frozen public-safe corpus
P1 evidence-path validation [PASS]
P2 full 120-case inference
post-inference oracle join
seam-specific correctness + calibration + reliability
P3 redacted public projection
independent publication verification [PASS]
published public evidence

Engineering principles

Disagreement is not improvement

A changed decision is only a candidate effect. Whether it is beneficial, harmful, or redundant is measured against independent ground truth.

Blinding protects the oracle

Ground-truth adjudicators cannot see deterministic or semantic outputs, probabilities, comparison results, or hidden construction tags.

Correctness is not operational policy

Raw semantic quality is measured separately from confidence thresholds, fallbacks, lifecycle rules, and the final applied decision.

Unresolved evidence stays unresolved

A genuine three-way conflict can remain explicitly unresolved; the study does not invent a label to preserve sample size.

Methodological shortcomings are evidence too

The authoritative v1 pass omitted structured justifications even though private reasoning summaries survived upstream. The frozen cohort is not rewritten; the corrective response is a versioned future contract with provider-neutral structured justification and end-to-end persistence checks.

Publication is a separate evidence contract

A valid private oracle was not automatically safe to publish. P3 uses a deliberately redacted public projection and an independent verifier for allowlist, integrity, cohort identity, and privacy boundaries.

Freeze executable surfaces, not just version labels

A runtime lock is only reproducible when it binds the result-affecting executable surface and rechecks it before qualification or execution. V2 hardening happened before new treatment outcomes, so it changes the readiness boundary rather than the result.

Publishability is outcome-neutral

The study succeeds when its evidence is defensible and reproducible, even if the semantic candidate is equal to or worse than deterministic control.

Routing Semantic v1 — published study result

audit 2026-09-27portfolio snapshot 2026-09-27
frozen audit

The completed study measures deterministic and semantic routing against independently frozen ground truth, then publishes a separately verified public projection. The result is heterogeneous by seam: semantic-candidate accuracy was higher for task class and interaction required, while semantic-risk ordinal error was approximately unchanged.

Frozen cases

120

public-safe shared corpus

Decision seams

3

task class · interaction required · semantic risk

Minimum eligible / seam

100

preregistered publication threshold

Paired bootstrap

10,000

deterministic resamples · 95% confidence

Oracle path

Frozen + validated

qualified DeepSeek adjudication → dispute-only C → recorded human resolution

Oracle API baseline

$0.07

36 DeepSeek-4.1-Flash requests · 726,263 tokens · through P0B freeze

Decision provenance

Corrected diagnosis

authoritative structured justification omitted · private reasoning summaries retained upstream

Task class

81.67% vs 52.50%

+29.17 pp · 95% +20.0 to +38.33

Interaction required

91.67% vs 71.67%

+20.0 pp · 95% +10.83 to +29.17

Semantic risk

0.28975 vs 0.29167 MAE

paired interval -0.101917 to +0.092917 · spans zero

Provider requests

120 / 120 successful

91,445 input + 10,598 output = 102,043 provider tokens

Request latency

p50 0.1468 s

p90 0.1865 s · p95 0.2206 s

Calibration

Brier + ECE / log loss

interaction 0.07741 Brier · 0.14942 ECE; task 0.27346 Brier · 1.80242 log loss; risk 0.32484 Brier · 0.53531 log loss

Decision changes

30/6 · 38/3

interaction beneficial/harmful · task beneficial/harmful; task also had 10 changed-but-both-wrong cases

P2 full cohort

COMPLETE

120 cases · 360 observations · 120 requests · 0 exclusions

P3 publication

VERIFIED + PUBLISHED

public oracle projection · integrity + privacy checks passed

AreaCurrent stateAssessmentNext action
Research questionAnswered for v1 cohortDifference was scored against independently frozen ground truth rather than treated as improvement by itselfReplicate/version future studies rather than generalize beyond the frozen cohort
Ground truthFrozen, validated, and joined post-inferenceIndependent and blinded before treatment inferenceKeep v1 immutable; any future adjudication change belongs to a new versioned study
Oracle baseline cost/usageRecorded$0.07 · 36 requests · 726,263 tokens via Codex harness -> codex-lb -> DeepSeek-4.1-FlashKeep this baseline separate from P1/P2 Jev treatment accounting
AdjudicationP0A/P0B completeDeepSeek-Flash cohort completed A/B/C plus recorded human resolutionDo not relabel the frozen oracle; preserve the completed cohort as baseline ground truth
Decision provenanceDiagnosis correctedAuthoritative A/B/C pass omitted structured justification; private Inspect reasoning summaries survived upstream but were unavailable to original human resolutionv2 requires provider-neutral structured justification plus round-trip persistence and explicit session-usage scope
Model vs policySeparatedRaw semantic correctness can be measured independently of thresholds/fallbackReport both layers without conflating them
P1 instrumentationPassedLive evidence path verified on 8 development-only cases with 0 exclusionsRetain as the preflight pattern for future cohorts
P2 outcomeCompleteHigher candidate accuracy on task class and interaction required; semantic-risk MAE approximately unchangedKeep reporting seam-specific; do not collapse into a universal winner
P2 provider evidenceComplete120/120 requests succeeded; 102,043 provider tokens; p50/p90/p95 request duration 0.1468/0.1865/0.2206 sKeep request-level efficiency separate from incomplete provider-cost evidence
Calibration edge caseCorrected + disclosedOne 0.99-mass vector in each multiclass seam; raw evidence preserved and derived calibration copy normalized onlyMove probability-mass validation into future preflight/reporting invariants
P3 publicationVerified + publishedPrivate-oracle disclosure path was replaced before release by a redacted public projection with independent verificationKeep private and public evidence contracts separate
Public claimsSeam-specific v1 findings publishedNo universal Jev superiority, downstream quality causality, cost advantage, or material semantic-risk improvement supportedRequire replication/new version for broader claims

Problem

  • The original question — whether TypeSafe/Jev disagrees with deterministic routing — measures behavioral difference, not correctness.
  • Successful provider calls, confidence values, and agreement rates cannot establish that a semantic decision is right without an independent reference.
  • A valid comparison also has to avoid leaking candidate/control outputs into ground-truth creation and must preserve disagreements or unresolved cases rather than quietly forcing a label.
  • Model quality and host policy are different questions: a semantic answer can be correct even when current thresholds reject it, or incorrect even when policy accepts it.

Technical Highlights

  • Frozen 120-case public-safe inference corpus with all three routing seams oracle-eligible
  • Independent oracle authoring view containing only case text, supplied metadata, eligibility, and the frozen decision rubric
  • Three routing seams: task class (Choice), interaction required (Noul), and semantic risk (Score)
  • A/B independent adjudication on the exact same hash-pinned blinded artifact
  • Independent C adjudication only for A/B-disputed cases and seams, without seeing A/B labels
  • Completed A/B/C oracle adjudication with recorded human resolution for genuine three-way conflicts
  • Oracle-baseline provider accounting through freeze: $0.07 USD, 36 DeepSeek-4.1-Flash API requests, 726,263 tokens
  • Oracle execution provenance: OpenAI Codex harness -> codex-lb load balancer -> DeepSeek-4.1-Flash API
  • Passed P1 development smoke: 8 cases, 8 provider requests, 24 per-seam observations, 0 exclusions
  • P1 verified oracle separation, privacy boundaries, probability persistence, unique request IDs, and single-count request-level accounting
  • Post-freeze evidence audit: the authoritative v1 A/B/C pass retained labels but not structured per-decision justifications; later private-log audit found provider reasoning summaries upstream, but they were not available to the original resolver and were not retrofitted into the frozen oracle
  • Completed P2 full study: 120 cases, 360 observations, 120 provider requests, 360 oracle outcomes, 0 exclusions
  • Published seam result: task class 81.67% semantic vs 52.50% control; interaction required 91.67% vs 71.67%; both paired 95% accuracy-difference intervals remained above zero
  • Semantic-risk result remained approximately unchanged: MAE 0.28975 semantic vs 0.29167 control, with the paired interval spanning zero
  • Change analysis: interaction required had 30 beneficial changes and 6 harmful changes; task class had 38 beneficial changes, 3 harmful changes, and 10 changed-but-both-wrong cases
  • Calibration: interaction-required Brier 0.07741 and ECE 0.14942; task-class multiclass Brier 0.27346 and log loss 1.80242; semantic-risk multiclass Brier 0.32484 and log loss 0.53531
  • P2 request evidence: 120/120 successful provider requests, 91,445 input tokens, 10,598 output tokens, 102,043 total provider tokens; p50 0.1468 s, p90 0.1865 s, p95 0.2206 s
  • P3 public projection passed exact-file allowlist, manifest-integrity, sample-identity, and privacy verification before publication
  • Lossless per-seam semantic evidence and single-count request-level latency/token/cost accounting
  • 95% confidence policy, deterministic 10,000-resample paired bootstrap, explicit percentile/ECE sample thresholds, and visible exclusion reason codes
  • Outcome-neutral publication policy: evidence quality is the gate, not a favorable semantic result
  • Routing Semantic v2 pre-live hardening binds and rechecks result-affecting executable surfaces including the host-side Jev implementation, exact TypeSafe SDK, Inspect/SWE runtime, Codex platform, Docker/Compose and resolved sandbox image identity, and the benchmark Inspect harness
  • V2 qualification preserves strict whole-completion parsing, schema/provenance identity, and the frozen v1 provenance shape separately from stronger scoped v2 provenance
  • Benchmark 0.6.3 produced the first full live IA-1 through IA-11 pass with sanitized ingress proof: 15 captured /v1/responses requests carried json_schema + strict=true with schema hashes matching persisted artifacts
  • Before any real label, real-cohort review measured the 120-case per-case schema at 225,971 serialized characters versus 3,310 for the eligibility-grouped equivalent — a 98.54% representation reduction with all 120 case identities preserved
  • Benchmark 0.6.4 records the grouped strategy as eligibility-grouped-case-enum/v1 and rejects the older qualification for real runs because the model-facing schema implementation changed
  • Current routing-semantic-v2 boundary: fresh 0.6.4 IA-1 through IA-11 qualification is required before real A/B; no real v2 A/B/C result is claimed
  • Agentic-Jev SWE-manager v1 completed 30 fresh paired tasks with 15/30 correct in both arms (0.0 pp observed difference; paired 95% bootstrap interval -16.7 to +20.0 pp); the treatment was Jev availability, not forced use
  • Jev was called on 20/30 v1 treatment attempts, but 0/20 successful calls met the frozen context-completeness diagnostic; exact provider-dispatch requests were preserved, localizing the loss to pre-dispatch request construction
  • The versioned manager v2 treatment uses deterministic request construction plus task-specific completeness policy, has passed live qualification on an already-observed task, and remains preregistered/unrun on its fresh effectiveness cohort

System Design

  • One frozen case is evaluated by deterministic control and one batched TypeSafe/Jev request projected into three decision observations.
  • The inference runner cannot receive oracle labels. The oracle is a separate artifact and is first joined after inference.
  • A and B label the same frozen authoring view independently. C receives only disputes. Genuine three-way conflicts require a recorded resolution or remain explicitly unresolved.
  • The completed oracle-baseline execution used the OpenAI Codex harness through codex-lb to DeepSeek-4.1-Flash; its cumulative provider/API accounting is tracked separately from P1/P2 TypeSafe/Jev traffic.
  • Raw TypeSafe evidence, Agent-Workflow's policy candidate, and the applied deterministic decision remain separate layers so model quality is not conflated with current operational thresholds or fallback rules.
  • Correctness and calibration are evaluated per seam; provider overhead is counted once per batched request rather than copied across three decision totals.
  • Before the full study, P1 exercised the same evidence path on 8 development-only cases and verified request/observation counts, probability persistence, oracle absence, privacy, runtime identities, and accounting invariants.
  • A future oracle cohort must emit compact structured decision justification in the same independent pass as each label and prove those fields survive persistence/export before real adjudication begins.
  • P2 reporting preserves raw probability evidence; when one vector in each multiclass seam summed to 0.99, only the derived calibration copy was normalized and the deviation was disclosed.
  • P3 uses a separate public oracle projection rather than treating the valid private frozen oracle as automatically publishable.
  • Benchmark tooling executes and sanitizes the experiment; comparative-eval owns comparison meaning; public results are published separately with hashes, identities, limitations, and manifests.
  • For v2, model-facing schema strategy is part of qualification identity: the 0.6.3 pass remains valid historical evidence, but 0.6.4's grouped real-cohort representation requires a fresh qualification before provider calls are admitted.

Deterministic Authority

  • Agent-Workflow remains application and lifecycle authority. Comparative mode keeps deterministic control authoritative while recording the semantic candidate as shadow/counterfactual evidence.
  • Comparative-eval does not schedule work, select models, mutate lifecycle state, or apply candidate decisions.
  • The oracle is independent evidence, not another source of operational authority.
  • Result-affecting changes to labels, rubrics, metrics, exclusions, or question meaning require a new study or dataset version.

Where Models Fit

  • TypeSafe/Jev supplies typed Choice, Noul, and Score evidence for the three registered routing seams.
  • Probability and confidence are evidence to measure and calibrate, not permission to perform host actions.
  • The study measures raw semantic correctness against the oracle separately from whether current Agent-Workflow policy would accept, reject, or fall back from that evidence.

Trust, Failure & Recovery

  • Oracle adjudicators must not see deterministic outputs, TypeSafe/Jev outputs, probability evidence, construction tags, comparison results, or other outside system knowledge excluded by the frozen authoring view.
  • Every case remains accounted for as eligible, failed, or excluded with a versioned reason; provider failures remain reliability evidence rather than disappearing from denominators.
  • If A/B disagree, C independently labels only the disputed case/seam. If A/B/C produce no majority, human resolution is recorded or the conflict remains explicitly unresolved.
  • The authoritative P0B pass preserves labels but not structured per-decision justifications. Human resolution therefore used the frozen case, supplied metadata, frozen rubric, and independent A/B/C labels. A later audit found provider reasoning summaries in private Inspect logs, but those summaries were supplementary execution evidence, were unavailable to the resolver, and are not being retrofitted into the frozen oracle.
  • The v1 oracle remains immutable as executed. The corrective design belongs in a subsequent protocol/contract version rather than a retroactive rerun presented as historical evidence.
  • Raw provider HTTP evidence and credentials remain private. P3 publishes a redacted oracle projection that retains final labels plus resolution status/method while removing adjudicator identities, individual A/B/C votes, pass hashes/timestamps, human-resolution rationale/participants, and private reasoning summaries.
  • The P3 verifier independently checks the exact publication allowlist, SHA manifest integrity, full frozen-sample identity, public-oracle projection contract, privacy flags, and absence of private adjudication/reasoning fields.
  • V2 qualification is fail-closed and representation-specific. A later implementation change to the real-cohort schema invalidates the earlier passing qualification for authorization purposes even though the earlier pass remains preserved as historical evidence.

Measured evidence

Frozen Routing Semantic v1 study contract

measured

120 frozen cases · 3 routing seams · 10,000 paired-bootstrap resamples

The preregistered study specification freezes the corpus target, three decision seams, sample thresholds, exclusion policy, oracle policy, metrics, and deterministic statistical policy before live comparative results are available.

Methodology and study-readiness evidence only. This does not establish that TypeSafe/Jev improves correctness, latency, cost, or any other outcome.

n=1202026-09-26
evidence method/source

Frozen independent oracle baseline

measured

P0A/P0B complete · DeepSeek-Flash adjudicator cohort · oracle frozen + validated

After IA-1 through IA-8 qualification, independent A/B passes labeled the exact blinded authoring view, C received only disputed seams, genuine three-way conflicts went through recorded human resolution, and the resulting oracle was frozen and validated against the 120-case corpus.

This establishes the independent ground-truth baseline later used for scoring. The oracle was frozen before treatment inference and joined only after the full P2 run completed.

n=1202026-09-26
evidence method/source

Oracle baseline API usage

measured

$0.07 USD · 36 requests · 726,263 tokens

Operator-recorded provider/API accounting for initial qualification and adjudication through P0B oracle freeze. Execution path: OpenAI Codex harness -> codex-lb load balancer -> DeepSeek-4.1-Flash API; study tooling recorded the routed alias openai-api/codex-lb/deepseek-flash.

Oracle-baseline construction only. Excludes P1 TypeSafe/Jev smoke, P2 comparative inference, and P3 publication. Raw billing/provider traffic is not committed, so these figures are not used to infer per-token pricing or comparative efficiency.

n=362026-09-26
evidence method/source

Oracle decision-provenance audit

measured

Authoritative labels retained · structured justifications omitted · private reasoning summaries retained upstream

Post-freeze evidence-chain review first identified that the labels-only authoritative pass omitted structured per-decision justification. Follow-up audit of retained private Inspect logs found contemporaneous provider reasoning-summary events, establishing that explanatory evidence existed upstream but was never promoted into the pass or original human-resolution workflow.

Methodology and evidence-completeness finding only. The private summaries are supplementary execution evidence, not hidden chain-of-thought, were not available to the original resolver, and are not used to rewrite the frozen v1 oracle.

n=1202026-09-27
evidence method/source

P2 full comparative result

measured

120 cases · 360 observations · 120 requests · 0 exclusions

All frozen cases ran with the oracle absent from inference. The frozen oracle was joined only after inference completed. Seam-specific correctness, calibration, disagreement, reliability, latency, and token evidence were then reported under the preregistered policy.

Bounded routing study. Task-class and interaction-required accuracy were higher for the semantic candidate in this cohort; semantic-risk MAE was approximately unchanged and its paired interval spans zero. No universal-superiority, downstream software-quality, or treatment-cost conclusion follows.

n=1202026-09-27
evidence method/source

P2 calibration evidence

measured

Interaction Brier 0.07741 · ECE 0.14942 · task-class Brier 0.27346 · risk Brier 0.32484

Calibration was computed from the persisted semantic probability evidence after the post-inference oracle join. One probability vector in each multiclass seam summed to 0.99; raw evidence was preserved and only the derived calibration copy was normalized.

Calibration evidence for the frozen routing-semantic-v1 cohort. Task-class log loss was 1.80242 and semantic-risk log loss was 0.53531. These metrics do not support a universal model-quality conclusion.

n=1202026-09-27
evidence method/source

P2 request-level provider evidence

measured

120/120 success · 102,043 tokens · p50 0.1468 s · p95 0.2206 s

Request-level accounting counted the single batched semantic provider request once per case rather than once per decision seam. The run recorded 91,445 input tokens, 10,598 output tokens, and 102,043 provider tokens total; latency was p50 0.1468 s, p90 0.1865 s, and p95 0.2206 s.

Provider request reliability, token, and request-duration evidence for P2 only. Provider cost evidence is incomplete, so these measurements do not support a treatment-cost conclusion.

n=1202026-09-27
evidence method/source

P3 verified public projection

measured

PASS · exact allowlist · manifest integrity · privacy contract

The publication verifier checked the exact file allowlist, SHA manifest, 120-case/360-observation cohort identity, redacted public oracle projection, privacy flags, and absence of private A/B/C votes, human-resolution rationale/identity, and private reasoning summaries.

Publication-integrity evidence. The public result is a deliberate projection of the valid private evidence bundle rather than a verbatim copy of the private frozen oracle.

n=1202026-09-27
evidence method/source

P1 live evidence-path smoke

measured

PASS · 8 cases · 8 requests · 24 observations · 0 exclusions

A deterministic development-only subset of the frozen corpus exercised the live TypeSafe/Jev path. Verification checked one request per observed case, three seam observations per case, unique request IDs, persisted semantic probabilities/distributions, explicit failures, oracle absence during inference, privacy-safe persistence, single-count request accounting, and frozen routing/v2 plus routing-state/v2 identities.

Instrumentation and evidence-integrity evidence only. The smoke is deliberately below the preregistered study threshold and does not support any TypeSafe/Jev comparative-effectiveness conclusion.

n=82026-09-26
evidence method/source

Routing Semantic v2 first full qualification pass

measured

PASS · IA-1 through IA-11 · 15 sanitized ingress requests

An archived Benchmark 0.6.3 retry completed the strengthened qualification. Sanitized loopback capture verified structured-output requests reached Codex-LB ingress with json_schema, strict=true, and schema hashes matching persisted primary/C artifacts; the qualification manifest reported qualified=true and all eleven gates passed.

Qualification evidence for the 0.6.3 schema implementation only. It does not authorize the later 0.6.4 grouped schema and is not a real-cohort oracle result.

n=152026-09-29
evidence method/source

Routing Semantic v2 real-schema scale review

measured

225,971 → 3,310 serialized chars · 98.54% reduction · 120/120 IDs preserved

Before any real label, deterministic reconstruction compared the per-case anyOf schema with an eligibility-grouped representation. All 120 frozen cases share the same three decision seams, so grouping removes repeated record shapes while retaining the exact allowed case-ID set and unchanged deterministic post-generation validation.

Representation and readiness evidence only. It does not establish that the larger schema would have failed at a provider, and it deliberately invalidates the 0.6.3 qualification for 0.6.4 authorization.

n=1202026-09-29
evidence method/source

Evidence & documentation