Things to make / Recipe 01 of 16 · Play
Ten puzzle levels, every one proven beatable
The agent invents a box-pushing puzzle, writes a solver, throws away every level it cannot beat, and shows you the proof inside the game.
- About 25 minutes
- For anyone who wants to see an agent check its own work instead of promising it.
- Runs in Claude Code and OMP
What the agent made
Ran it for real with Claude Code 2.1.295 and OMP 18.8.6 on 2026-10-09, in an empty session with no saved memory. Your result will differ.
Open what Claude Code made, full screen
The same prompt in OMP made its own version: open what OMP made. Two runs, two different results: that is normal.
Make your own
Get the files
Make an empty folder called
cli-practicein Downloads. No downloads: start in an empty folder.Open your agent in that folder
Open a terminal in
cli-practiceand typeclaudeoromp --approval-mode always-ask. First time? Set up in about five minutes, with steps for Mac, Windows, and Linux.Paste this prompt
Make a box-pushing puzzle game with ten levels, each one proven beatable. Work in this empty practice folder. 1. Write solver.py (Python 3 standard library only) that finds the shortest solution, in moves, for a level by breadth-first search over player and box positions. 2. Generate or design many candidate levels, run the solver on each, throw away every level it cannot beat, and keep ten that get harder, using each level's shortest solution length as its par. Save the ten levels with their solutions in levels.json. 3. Make python3 solver.py levels.json re-solve every level from scratch, confirm each stored solution wins in exactly par moves, print one line per level, and exit 0 only if all ten pass. Run it and report its exit status. 4. Build one self-contained puzzle.html that embeds the same ten levels: arrow keys and tappable on-screen buttons, undo, restart, a move counter against par, and a Watch the proof button that replays the stored solution one move at a time through the game's own move function, so the replay obeys the same rules a player does. Print the SHA-256 of levels.json in the page footer and in your final reply. Work only inside this folder; never install packages, use the network, call APIs, or load external assets. The page must work offline and on a phone. Use no em dashes or en dashes.
The prompt asks the agent to stay in this folder; it cannot enforce that. Read each file change and command before you approve it.
Check it yourself
Open
cli-practice/puzzle.htmlin your web browser (double-click it), then:Human check: Run python3 solver.py levels.json and see ten passing lines and exit 0. Run shasum -a 256 levels.json and match it to the page footer. Beat level 1 in par, then open level 10, press Watch the proof, and watch it win.
When we checked the recording above: python3 solver.py levels.json re-solved all ten levels and exited 0; the page footer hash matches shasum of levels.json (07a9a9b2...). Played level 1 with the arrow keys and finished in par, 15 moves. On level 10, Watch the proof replayed 99 moves through the game and ended: Proof checked: solved in exactly 99 moves, the par. Tapping the on-screen buttons moves the player. No console errors on a phone-size screen.
Do it for real
Use the same loop on anything you make with an agent: ask it for a checker first, then the thing, then make it run the checker and show you the result. Never accept done without the checker's output.
Why an agent, not a chat window
A chat reply can describe a puzzle. An agent can write a solver, run it on dozens of candidate levels, keep only the ones it proved, and replay each proof through the game's own rules on your screen.