Field note /
The first fresh manager study found zero observed lift — and a treatment flaw
Across 30 fresh paired SWE-Lancer manager tasks, control and treatment both scored 15/30; the request audit then found that 0/20 Jev calls met the frozen context-completeness threshold.
How the claim changed
Belief
After removing the earlier self-gating failure, a fresh paired manager cohort could estimate whether making live Jev evidence available changed proposal-selection correctness.
Evidence
All 30 expected pairs completed. Control accuracy was 15/30 and treatment accuracy was 15/30, for an observed 0.0 percentage-point difference. The paired 95% bootstrap interval was -16.7 to +20.0 points and exact McNemar p was 1.0.
Contradiction
The neutral aggregate was not the whole result. Jev was called on 20/30 treatment attempts and often influenced the final selection, yet 0/20 successful calls met the preregistered context-completeness diagnostic.
Correction
Preserve the v1 result as evidence for the treatment that actually ran. Do not repair it in place. Treat request construction as part of the experimental treatment and move authoritative context projection into deterministic code for v2.
Current claim boundary
The v1 result does not support a claim of improved accuracy, but the sample is too small to rule out meaningful positive or negative effects. Called-only subsets are self-selected and descriptive, not causal estimates.