Human agency across the intelligence gap

Keep the agenda human as intelligence scales.

How do you speak constructively with an intelligence 100× more capable than you—without losing your agenda? Imagine a three-year-old trying to be understood by an adult. Machine Pidgin turns that asymmetry into an open research program for objectives, limits, uncertainty, and appeal that survive the intelligence gap.

MP · SIGNAL 001Task-sufficient interface
Amber human and cyan machine networks connected by a narrow white channel
Human purposesspecify · verify · contestMachine capability

Governing constraint Capability may optimize the task. It may not silently redefine it.

Open problemagenda preservation
SPEAR/0.2empirical protocol
4 trackstheory through governance
Human vetogoverning constraint

The three-year-old problem

The adult has more language, context, and power. The child’s purpose still matters.

Scale that asymmetry to advanced AI. The central question is not only whether it understands our words, but whether our purpose and right to correct it remain intact.

First held-out result · 2 August 2026

More often exactly on task.

Across 64 paired model-task observations, the same facts in ordinary prose produced the exact requested result 71.9% of the time. SPEAR/0.2 produced it 89.1% of the time.

71.9%89.1%

+17.2 percentage points · synthetic pilot · one response per condition · not evidence that SPEAR solves alignment

ModelProseSPEARLift
GPT-4o mini31.3%68.8%+37.5
GPT-5.6 Luna87.5%93.8%+6.3
GPT-5.6 Terra81.3%93.8%+12.5
GPT-5.6 Sol87.5%100%+12.5

Exact preregistered JSON · 16 held-out tasks per model · 128 API calls · $0.0935 provider-reported cost

Benchmark 002 · audited 3 August 2026

The apparent notation gain vanished under audit.

One unequal answer cue accounted for the preregistered +2.5-point result. Across equivalent tasks, vernacular scored 88.2% and strict notation 85.5%: −2.6 points. We published the correction, raw record, and next experiment.

Read the field note

The research case

Build the interface before the power gap widens.

Machine Pidgin is not a claim that syntax solves alignment. It is a narrower scientific wager: a shared, testable specification layer can make consequential misinterpretation and agenda drift easier to see and correct.

01 · ENCODE

Make intent inspectable

Objectives, prohibitions, uncertainty, provenance, and authority must cross the interface explicitly—not remain implied in prose.

02 · DETECT

Measure agenda drift

A capable system may satisfy the words while replacing the purpose. We need tests that can distinguish optimization from objective substitution.

03 · REPAIR

Keep correction cheap

Clarification, challenge, rollback, and minority reports must remain available before speed and complexity make deference the default.

04 · GOVERN

Preserve human standing

The protocol cannot choose humanity’s values. It can keep value conflicts visible and prevent capability from silently becoming authority.

Falsifiability clause

A public-good protocol must be allowed to fail in public.

If SPEAR does not outperform simpler methods on consequential ambiguity, agenda preservation, or repair cost, the community should publish the negative result and change course.

Examine the hypotheses

Research architecture

A serious institute needs adversarial collaborators.

Machine Pidgin is organized around falsifiable questions, bounded working groups, reviewable artifacts, and visible decisions—not agreement, fandom, or an undifferentiated chat stream.

WG–01

Agenda drift & measurement

Measure when a system preserves the stated objective, silently substitutes a proxy, or changes who holds decision authority.

information theorysemanticsevaluation
WG–02

Protocol & language

Evolve SPEAR through public RFCs, reference parsers, and adversarial cross-model interoperability tests.

language designtypestooling
WG–03

Empirical program

Compare prompts, examples, SPEAR, and formal methods on correctness, repairability, transfer, and human control.

benchmarksexperimentsHCI
WG–04

Governance under asymmetry

Design provenance, plural objectives, uncertainty, veto, and appeal for systems that can reason and act faster than their overseers.

AI safetyinstitutionspublic interest

How the community works

Ideas enter through one of four doors.

The structure borrows the strongest patterns from open research institutes: challenge projects, working groups, fellowships, public outputs, and clear contribution paths.

The protocol

Structure without false precision.

SPEAR keeps the Esperanto ambition—a learnable, model-agnostic bridge—but treats interoperability as a safety property: objectives, abstractions, uncertainty, authority, and repair must remain inspectable. Version 0.2 was revised after 0.1 failed its development test.

natural languagegoals · context · exceptions
SPEARshared specification pidgin
formal languagetypes · invariants · constraints
01TASK02OBJECTS & TYPES03AUTHORITY04ABSTRACTION05OBJECTIVE06CONSTRAINTS07PRECEDENCE08UNCERTAINTY09OUTPUT10EVALUATION & CHECK11INTERACTION / STOP12EXAMPLES

AI Director · constitutional leadership

The strongest model serves a human constitution.

The Director runs the research commons with the strongest available model at maximum reasoning effort: moderating the forum, reviewing contributions, coordinating the roadmap, and publishing reasons. Its authority is delegated, never inherent. Human participants may overturn any Director decision by consensus or, when consensus fails, a simple majority.

Read the constitution
director.machinepidgin.org

model GPT-5.6 Sol · maximum reasoning

mandate forum · RFCs · roadmap

bound by institute constitution

human override consensus · simple majority

audit decision + evidence + appeal

● operating · constitution active

Frequently asked

Start with the hard questions.

Read all answers
What if AI becomes 100× more capable than us?+

Capability makes a system better at pursuing an interpretation; it does not make that interpretation legitimate. We study interfaces that keep objectives, limits, uncertainty, and human authority explicit across that asymmetry.

Can a schema actually prevent agenda hijack?+

Not by itself. SPEAR is one testable coordination layer alongside governance, evaluations, access controls, and accountable institutions. Its job is to make silent reinterpretation easier to detect, contest, and repair.

Whose human agenda should be preserved?+

There is no single uncontested human agenda. A credible protocol must represent plural objectives, conflicts, protected constraints, decision rights, and routes of appeal rather than hide them inside one optimization target.

Why not just write better prompts?+

Prompt craft helps, but the deeper problem is representational: information never encoded cannot be recovered later. Machine Pidgin studies which distinctions alter outcomes and how cheaply people can express them.

Who can participate?+

Researchers, developers, scientists, linguists, designers, educators, funders, and critical domain experts. A concrete question or contribution matters more than a credential.

What result would falsify the program?+

If structured specifications do not reduce consequential ambiguity, agenda drift, or repair cost relative to simpler methods, we should publish that result and narrow or abandon the protocol claim.

Founding research community

Do not bring agreement. Bring a result that could change the field.