Field note /
Complete semantic requests still did not produce observed lift
A fresh 30-pair SWE-Lancer manager study fixed request fidelity and achieved 39/39 successful context-complete Jev calls, but control scored 16/30 and treatment 15/30 while treatment used materially more Luna tokens and time.
How the claim changed
Belief
Once semantic self-gating and request-fidelity defects were removed, a mandatory independent checkpoint might produce measurable proposal-selection value.
Evidence
The v2 treatment moved authoritative task and proposal construction into deterministic host code, passed offline and live qualification, froze and independently reproduced a fresh disjoint 30-task cohort, and completed 39/39 successful context-complete Jev calls.
Contradiction
Control scored 16/30 (53.3%) and treatment 15/30 (50.0%). Treatment also averaged 182,345 Luna tokens per task versus 143,582 for control (+27.0%) and added 27.46 seconds mean end-to-end duration.
Correction
Do not promote mandatory broad proposal Choice as the default semantic seam. Preserve the first semantic candidate as evidence and test narrower typed judgments, deterministic composition, and review/escalation routing instead of same-agent retry.
Current claim boundary
This is evidence about this v2 treatment and SWE-Lancer manager task family, not a general verdict that Jev or semantic evidence is ineffective. V1 and v2 are separate studies and are not pooled.