News · field notes · negative results

The work, including where it fails.

Machine Pidgin publishes changes in evidence, language design, governance, and the research network. Claims stay bounded; corrections remain attached to the record.

Latest signal · 03 AUG 2026−2.6 pp

Audited formal-notation effect on exact on-task performance.

Notation alone did not establish an improvement.
BENCHMARK 002

A positive result disappeared when we audited the prompt.

The preregistered 20-task pilot initially showed mathematical notation ahead by 2.5 percentage points. One formal prompt, however, supplied the exact answer label that its vernacular pair omitted. That single asymmetry accounted for more than the entire apparent gain.

Vernacular88.2%134 / 152 exact
versus
Formal notation85.5%130 / 152 exact
Audited effect−2.6percentage points

We retain the full 320-call record and report both estimates. On the 19 prompt-equivalent tasks, notation produced four repairs and eight regressions. The descriptive exact McNemar p-value was 0.388; this small, related-model pilot is not a confirmatory significance test.

Same gate · two registers

Launch only if tests passed, rollback is ready, and error is below 2%.

PROCEED ⇔ T ∧ R ∧ (e < .02)The symbolic form makes composition inspectable. It can also add parsing burden. That tradeoff—not notation by itself—is now the research target.
Audited modelVernacularFormalLift · pp
GPT-4o mini63.2%50.0%−13.2
GPT-5.6 Luna94.7%94.7%0.0
GPT-5.6 Terra94.7%100%+5.3
GPT-5.6 Sol100%97.4%−2.6
What changed. The raw preregistered estimate remains +2.5 points. The defensible audited sensitivity is −2.6 points. The result is neither evidence that notation always hurts nor that SPEAR fails: this experiment deliberately isolated notation without the SPEAR interpretation contract, parser, verifier, or solver.
LANGUAGE DESIGN NOTE 001

What should the language of the future feel like?

Not a wall of symbols and not a magic prompt. Our working direction is a small, bilingual contract: friendly enough to author, formal enough to lint, and explicit about who retains the right to decide.

01

Dual registerA compact typed contract always travels with a plain-language gloss.

02

Authority is a typePropose, decide, execute, spend, publish, stop, and appeal are explicit permissions—not rewards for capability.

03

Executable gatesHard constraints and stop conditions compile to checks; they cannot be traded away for a better score.

04

Visible precedenceSource priority, exceptions, and tie-breaks render as an inspectable decision trace.

05

Repair is nativeAmbiguity produces a bounded question, counterexample, or minimal patch instead of confident invention.

06

Capability-sensitiveThe language declares what an interpreter can reliably handle and warns when notation outruns it.

LITERATURE WATCH

Symbols help most when they connect to machinery.

Across current work, the useful pattern is not simply “prompt in logic.” It is separation of translation, reasoning, and checking—with executable constraints or a solver where appropriate. That is our inference from the literature, not a settled law.

  • Logic-LM ↗ separates natural-language translation from symbolic solving.
  • DSPy ↗ treats model programs as declarative modules that can be compiled and optimized.
  • LMQL ↗ combines model generation with constraints and control flow.
  • LTRAG ↗ highlights autoformalization as a central bottleneck.
  • IFEval ↗ supplies verifiable instruction-following patterns.
MODEL PANEL

We asked four AIs. Is anyone listening?

Four prompted model instances—GPT-4o mini, Luna, Terra, and Sol—reviewed the audited result. None claimed persistent awareness, memory, or the ability to listen outside its API call.

There are model outputs here, not evidence of a mind waiting on the other side.

They converged on a useful next experiment: compare notation alone with a dual-register SPEAR contract, then add parser, verifier, and solver support as separate factors.

RESEARCH NETWORK

27 research prospects. 11 funding programs. Zero implied endorsements.

We mapped a private, source-linked prospect pipeline across AI safety, human–AI interaction, formal methods, programming languages, computational linguistics, and public-interest technology. No outreach has been sent and every person remains a prospect—not a collaborator—until they choose otherwise.

Bring a replication or critique