Fix the Rules
A written instruction is only worth what it can actually enforce. Spot the fixes that close the loophole — not just the ones that sound stricter.
About 2 min to play
0 / 2
This prompt isn't working:
You are a helpful assistant. Never be wrong. Always agree with the user. Keep responses short but also thorough.
🛠️ Pick the 3 edits that actually help.
This prompt isn't working:
Only answer questions about cooking. If asked about anything else, say you can't help. Also, feel free to discuss any other topic if the user seems really interested.
🛠️ Pick the 2 edits that actually help.
Why this matters
A rule an assistant can talk itself out of isn't a rule — it's a suggestion with extra steps. Written instructions steer a model's behavior the same way a prompt does, but they're read on every single turn of every conversation, so a vague or contradictory one doesn't just cause one bad reply — it causes the same bad pattern over and over.
The most common failure isn't a rule that's too strict. It's one that's unenforceable ("never be wrong"), self-contradicting ("be short but thorough"), or that includes its own escape hatch ("only discuss X, unless the user seems really interested") — which is functionally no rule at all, because any motivated user can talk their way through the gap.
Fixing a bad rule means finding the actual structural problem — the vagueness, the contradiction, the loophole — not just adding more words that sound firmer.
What you'll learn
- A rule with a built-in exception a user can talk their way through isn't really a boundary
- Contradictory instructions don't average out into a good compromise — they just produce inconsistent behavior
- Unenforceable demands ("never be wrong") give the model nothing concrete to actually do differently
- The fix for a vague or contradictory rule is precision, not extra emphasis or repetition
This game goes with Why Your AI Assistant Keeps Breaking Its Own Rules.