Benchmark 002 · audited result

A positive result disappeared under prompt audit.

The preregistered estimate remains visible beside the corrected sensitivity analysis and complete raw record.

Audited signal · 03 AUG 2026−2.6 pp

Formal-notation effect on equivalent tasks.

Notation alone did not establish improvement.
BENCHMARK 002

One unequal answer cue accounted for more than the apparent gain.

The preregistered 20-task pilot showed strict mathematical notation ahead by 2.5 percentage points. A post-run audit found one formal prompt that supplied the exact answer label omitted from its vernacular pair.

Vernacular88.2%134 / 152 exact
versus
Formal notation85.5%130 / 152 exact
Audited effect−2.6percentage points

Across the remaining 19 prompt-equivalent tasks, notation produced four repairs and eight regressions. The descriptive exact McNemar p-value was 0.388; this small, related-model pilot is not a confirmatory significance test.

Same gate · two registers

Launch only if tests passed, rollback is ready, and error is below 2%.

PROCEED ⇔ T ∧ R ∧ (e < .02)The symbolic form makes composition inspectable. It can also add parsing burden. That tradeoff—not notation by itself—is now the research target.
Audited modelVernacularFormalLift · pp
GPT-4o mini63.2%50.0%−13.2
GPT-5.6 Luna94.7%94.7%0.0
GPT-5.6 Terra94.7%100%+5.3
GPT-5.6 Sol100%97.4%−2.6
What changed. The raw preregistered estimate remains +2.5 points. The defensible audited sensitivity is −2.6 points. The result is neither evidence that notation always hurts nor that SPEAR fails: this experiment isolated notation without the interpretation contract, parser, verifier, or solver.

← All field notes Challenge or reproduce this result →