More exact answers, under a deliberately narrow test.
Sixteen held-out synthetic tasks were run in ordinary prose and SPEAR/0.2 across four model tiers. Exact normalized adherence rose from 46/64 observations (71.9%) to 57/64 (89.1%).
The tasks tested constrained selection, scheduling, authority, precedence, exceptions, information value, stop gates, and exact JSON contracts. All 128 held-out responses were valid JSON. Provider-reported held-out cost was $0.0935, and the SPEAR condition used more prompt tokens.
The development audit is part of the result. An impossible answer key, inconsistent canonical label, arithmetic errors, and a provider rate limit were retained and documented before the held-out run. The evidence supports further independent replication—not standardization.
