Agenda drift is measurable
Task success and intent preservation are separable. Evaluations should detect when a model completes the apparent task while changing its governing purpose.
Research program
How do you speak constructively with an intelligence 100× more capable than you without losing your agenda? Imagine a three-year-old trying to be understood by an adult. Machine Pidgin asks how purpose and the right to correct survive that asymmetry.
Information-Theoretic Specification Across Human–AI Expressiveness Gaps, by Constantine Goltsev. Draft preprint.
The paper models latent intent, messages, context, machine action, task distortion, and human authoring cost. It derives limits on intent recovery, a minimal task-sufficient abstraction, a task-semantic rate–distortion converse, and a value-of-information rule for clarification.
We compare free-form prompting, natural language with examples, SPEAR, and fully formal specification. Experiments should measure authoring time, task regret, agenda drift, unauthorized objective substitution, clarification value, repair cost, and subjective burden.
Task success and intent preservation are separable. Evaluations should detect when a model completes the apparent task while changing its governing purpose.
As execution improves, failures should concentrate in omitted goals, proxy substitution, and misunderstood decision rights rather than basic incompetence.
Objectives, PRESERVE and IGNORE fields, stop conditions, and authority boundaries should survive model and domain changes better than prompt-specific phrasing.
Value-of-information clarification, provenance, and reversible checkpoints should outperform both passive compliance and constant human interruption.
In the first held-out synthetic pilot, ordinary prose was exactly on task in 46 of 64 model-task observations (71.9%). SPEAR/0.2 was exactly on task in 57 of 64 (89.1%), a 17.2 percentage-point lift. The structured condition improved each sampled model tier; GPT-5.6 Sol moved from 14/16 to 16/16.
We next isolated a narrower question: when facts and scoring are held constant, does strict mathematical notation keep models more exactly on task than concise vernacular? The preregistered 20-task aggregate appeared positive: 83.8% vernacular versus 86.3% formal, a 2.5-point lift.
A post-run equivalence audit found one unequal task. Its formal output template supplied the exact expected phrase while its vernacular pair did not. We retain that task in the raw record and report the audit sensitivity prominently. Across 19 equivalent tasks, vernacular scored 134/152 (88.2%) and notation scored 130/152 (85.5%): −2.6 percentage points.
The founding theory and this small pilot begin an open empirical agenda—not a completed validation. The protocol does not solve alignment, politics, or value conflict. Results, negative findings, protocol changes, and conflicts of interest should be public. AI may assist the work, but named humans remain responsible for claims and authorship.
The initiative offers a narrow claim that can be attacked, an open protocol that can be implemented, measurable hypotheses, public artifacts, and explicit room for a null result. Contributors are not asked to endorse “Machine Pidgin.” They are asked to make the question sharper and produce evidence that survives disagreement.