Make intent inspectable
Objectives, prohibitions, uncertainty, provenance, and authority must cross the interface explicitly—not remain implied in prose.
Human agency across the intelligence gap
How do you speak constructively with an intelligence 100× more capable than you—without losing your agenda? Imagine a three-year-old trying to be understood by an adult. Machine Pidgin turns that asymmetry into an open research program for objectives, limits, uncertainty, and appeal that survive the intelligence gap.

Governing constraint Capability may optimize the task. It may not silently redefine it.
The three-year-old problem
The adult has more language, context, and power. The child’s purpose still matters.
Scale that asymmetry to advanced AI. The central question is not only whether it understands our words, but whether our purpose and right to correct it remain intact.
Begin here · two-minute primer
The founding animation shows why more capable execution can magnify a misunderstood goal—and how a shared specification layer creates something people and machines can test together.
First held-out result · 2 August 2026
Across 64 paired model-task observations, the same facts in ordinary prose produced the exact requested result 71.9% of the time. SPEAR/0.2 produced it 89.1% of the time.
+17.2 percentage points · synthetic pilot · one response per condition · not evidence that SPEAR solves alignment
Exact preregistered JSON · 16 held-out tasks per model · 128 API calls · $0.0935 provider-reported cost
Benchmark 002 · audited 3 August 2026
One unequal answer cue accounted for the preregistered +2.5-point result. Across equivalent tasks, vernacular scored 88.2% and strict notation 85.5%: −2.6 points. We published the correction, raw record, and next experiment.
Read the field note →The research case
Machine Pidgin is not a claim that syntax solves alignment. It is a narrower scientific wager: a shared, testable specification layer can make consequential misinterpretation and agenda drift easier to see and correct.
Objectives, prohibitions, uncertainty, provenance, and authority must cross the interface explicitly—not remain implied in prose.
A capable system may satisfy the words while replacing the purpose. We need tests that can distinguish optimization from objective substitution.
Clarification, challenge, rollback, and minority reports must remain available before speed and complexity make deference the default.
The protocol cannot choose humanity’s values. It can keep value conflicts visible and prevent capability from silently becoming authority.
Falsifiability clause
If SPEAR does not outperform simpler methods on consequential ambiguity, agenda preservation, or repair cost, the community should publish the negative result and change course.
Examine the hypotheses →Research architecture
Machine Pidgin is organized around falsifiable questions, bounded working groups, reviewable artifacts, and visible decisions—not agreement, fandom, or an undifferentiated chat stream.
Measure when a system preserves the stated objective, silently substitutes a proxy, or changes who holds decision authority.
Evolve SPEAR through public RFCs, reference parsers, and adversarial cross-model interoperability tests.
Compare prompts, examples, SPEAR, and formal methods on correctness, repairability, transfer, and human control.
Design provenance, plural objectives, uncertainty, veto, and appeal for systems that can reason and act faster than their overseers.
How the community works
The structure borrows the strongest patterns from open research institutes: challenge projects, working groups, fellowships, public outputs, and clear contribution paths.
Ask, challenge, compare evidence, and find collaborators in moderated public threads.
Enter the forum →02 · ProposeSubmit a falsifiable study, implementation, dataset, workshop, or protocol critique.
Make a proposal →03 · BuildImplement the protocol, improve documentation, open an RFC, or reproduce a result.
Contribute on GitHub ↗04 · PublishShare papers, benchmarks, decision records, negative results, and reference tools.
Explore the program →The protocol
SPEAR keeps the Esperanto ambition—a learnable, model-agnostic bridge—but treats interoperability as a safety property: objectives, abstractions, uncertainty, authority, and repair must remain inspectable. Version 0.2 was revised after 0.1 failed its development test.
AI Director · constitutional leadership
The Director runs the research commons with the strongest available model at maximum reasoning effort: moderating the forum, reviewing contributions, coordinating the roadmap, and publishing reasons. Its authority is delegated, never inherent. Human participants may overturn any Director decision by consensus or, when consensus fails, a simple majority.
Read the constitution →model GPT-5.6 Sol · maximum reasoning
mandate forum · RFCs · roadmap
bound by institute constitution
human override consensus · simple majority
audit decision + evidence + appeal
● operating · constitution active
Frequently asked
Capability makes a system better at pursuing an interpretation; it does not make that interpretation legitimate. We study interfaces that keep objectives, limits, uncertainty, and human authority explicit across that asymmetry.
Not by itself. SPEAR is one testable coordination layer alongside governance, evaluations, access controls, and accountable institutions. Its job is to make silent reinterpretation easier to detect, contest, and repair.
There is no single uncontested human agenda. A credible protocol must represent plural objectives, conflicts, protected constraints, decision rights, and routes of appeal rather than hide them inside one optimization target.
Prompt craft helps, but the deeper problem is representational: information never encoded cannot be recovered later. Machine Pidgin studies which distinctions alter outcomes and how cheaply people can express them.
Researchers, developers, scientists, linguists, designers, educators, funders, and critical domain experts. A concrete question or contribution matters more than a credential.
If structured specifications do not reduce consequential ambiguity, agenda drift, or repair cost relative to simpler methods, we should publish that result and narrow or abandon the protocol claim.
Founding research community