Residual entropy beats prompt length
Errors should track task-relevant ambiguity more closely than the number of tokens written.
Research program
Machine Pidgin treats human–AI specification as a semantic communication problem: transmit the smallest abstraction that preserves the distinctions a task actually needs.
Information-Theoretic Specification Across Human–AI Expressiveness Gaps, by Constantine Goltsev. Draft preprint.
The paper models latent intent, messages, context, machine action, task distortion, and human authoring cost. It derives limits on intent recovery, a minimal task-sufficient abstraction, a task-semantic rate–distortion converse, and a value-of-information rule for clarification.
We will compare free-form prompting, natural language with examples, SPEAR, and fully formal specification. Experiments should measure authoring time, authoring error, task regret, clarification count, and subjective burden.
Errors should track task-relevant ambiguity more closely than the number of tokens written.
Too little detail raises ambiguity; too much raises human cost and creates false precision.
Explicit PRESERVE and IGNORE fields should generalize better across models and domains.
Value-of-information clarification should outperform both never asking and always asking.
This is a theoretical position and an open empirical agenda—not a completed validation. Results, negative findings, protocol changes, and conflicts of interest should be public. AI may assist the work, but named humans remain responsible for claims and authorship.