Skip to content
Nate G. / Vancouver, WA · US

Field note /

The first fresh manager study found zero observed lift — and a treatment flaw

Across 30 fresh paired SWE-Lancer manager tasks, control and treatment both scored 15/30; the request audit then found that 0/20 Jev calls met the frozen context-completeness threshold.

Study ResultEvaluation · SWE-Lancer manager decisions30-pair v1 completed · sanitized public result bundle pending

How the claim changed

Belief

After removing the earlier self-gating failure, a fresh paired manager cohort could estimate whether making live Jev evidence available changed proposal-selection correctness.

Evidence

All 30 expected pairs completed. Control accuracy was 15/30 and treatment accuracy was 15/30, for an observed 0.0 percentage-point difference. The paired 95% bootstrap interval was -16.7 to +20.0 points and exact McNemar p was 1.0.

Contradiction

The neutral aggregate was not the whole result. Jev was called on 20/30 treatment attempts and often influenced the final selection, yet 0/20 successful calls met the preregistered context-completeness diagnostic.

Correction

Preserve the v1 result as evidence for the treatment that actually ran. Do not repair it in place. Treat request construction as part of the experimental treatment and move authoritative context projection into deterministic code for v2.

Current claim boundary

The v1 result does not support a claim of improved accuracy, but the sample is too small to rule out meaningful positive or negative effects. Called-only subsets are self-selected and descriptive, not causal estimates.

Public evidence