One unequal answer cue accounted for more than the apparent gain.
The preregistered 20-task pilot showed strict mathematical notation ahead by 2.5 percentage points. A post-run audit found one formal prompt that supplied the exact answer label omitted from its vernacular pair.
Vernacular88.2%134 / 152 exact
versusFormal notation85.5%130 / 152 exact
Audited effect−2.6percentage points
Across the remaining 19 prompt-equivalent tasks, notation produced four repairs and eight regressions. The descriptive exact McNemar p-value was 0.388; this small, related-model pilot is not a confirmatory significance test.
Same gate · two registers
Launch only if tests passed, rollback is ready, and error is below 2%.
PROCEED ⇔ T ∧ R ∧ (e < .02)The symbolic form makes composition inspectable. It can also add parsing burden. That tradeoff—not notation by itself—is now the research target.Audited modelVernacularFormalLift · pp
GPT-4o mini63.2%50.0%−13.2
GPT-5.6 Luna94.7%94.7%0.0
GPT-5.6 Terra94.7%100%+5.3
GPT-5.6 Sol100%97.4%−2.6
What changed. The raw preregistered estimate remains +2.5 points. The defensible audited sensitivity is −2.6 points. The result is neither evidence that notation always hurts nor that SPEAR fails: this experiment isolated notation without the interpretation contract, parser, verifier, or solver.
