Behavior
Examples prove named outcomes. State-machine traces exercise lifecycles. Property and fuzz cases cover broad input spaces without prescribing an implementation.
PromptGolf asks one question: did the submitted spec cause an agent to build the requested product capabilities? Every point must trace to positive, observable evidence.
Examples prove named outcomes. State-machine traces exercise lifecycles. Property and fuzz cases cover broad input spaces without prescribing an implementation.
A requirement tree connects every scored capability to the public task, a domain rule, and the evidence needed to satisfy it.
Framework-aware adapters build and start the submitted workspace, then expose its routes, controls, outputs, and commands through one canonical protocol.
Players can inspect the requirement tree, evidence methods, capability vocabulary, scoring weights, and run results. Hidden evaluator source and test cases remain outside the builder workspace so the artifact must satisfy the capability rather than memorize the judge.
PromptGolf does not use mutation or negative testing, implementation resemblance, source layout, fixed signatures where valid alternatives exist, CSS fingerprints, preferred methods, or model-specific prompt wording. Invalid and failure-state behavior can still earn positive evidence when the brief requires that observable capability.