Write the spec; the agent builds it; the hidden tests grade what you shipped. Each hole is a normal-looking public brief whose hidden tests grade the messy reality a serious spec encodes: cents math, race conditions, the empty state nobody specs. Fewer prompts, lower score.
These put you in a loop instead of grading one prompt: route a team of agents to ship a task, tune a retrieval pipeline until it surfaces buried answers, or build the evaluator that catches a model’s mistakes.
A checkout that's quietly wrong. Find the store's hidden rules by running it.
Don't write the checkout. Direct a team of agents to ship it right.
Every other hole grades your spec. This one grades how you drive the build.
Don't fix one build. Author one orchestration that survives a world it never saw.
Two domains, one number. Route the symptom to the station that owns it.
Build the retrieval, not the answer. Tune a search pipeline until it finds what's buried.
You don't fix the AI. You build the eval that catches when it's wrong.