Benchmark 001 · held-out pilot

SPEAR improved exact adherence in a small synthetic pilot.

A permanent field note with the exact claim, limits, costs, and reproducible record.

Held-out signal · 02 AUG 2026+17.2 pp

Exact on-task adherence versus ordinary prose.

Bounded pilot · not peer reviewed
RESEARCH NOTE 001

More exact answers, under a deliberately narrow test.

Sixteen held-out synthetic tasks were run in ordinary prose and SPEAR/0.2 across four model tiers. Exact normalized adherence rose from 46/64 observations (71.9%) to 57/64 (89.1%).

Ordinary prose71.9%46 / 64 exact
versus
SPEAR/0.289.1%57 / 64 exact
Observed effect+17.2percentage points

The tasks tested constrained selection, scheduling, authority, precedence, exceptions, information value, stop gates, and exact JSON contracts. All 128 held-out responses were valid JSON. Provider-reported held-out cost was $0.0935, and the SPEAR condition used more prompt tokens.

ModelProseSPEARLift · pp
GPT-4o mini31.3%68.8%+37.5
GPT-5.6 Luna87.5%93.8%+6.3
GPT-5.6 Terra81.3%93.8%+12.5
GPT-5.6 Sol87.5%100.0%+12.5
Claim boundary. This is a small synthetic pilot designed by the protocol author, with related model families and strict exact-output scoring. It does not show that SPEAR recovers unexpressed values, prevents deception, resolves political disagreement, or solves alignment.

The development audit is part of the result. An impossible answer key, inconsistent canonical label, arithmetic errors, and a provider rate limit were retained and documented before the held-out run. The evidence supports further independent replication—not standardization.

← All field notes Challenge or reproduce this result →