Totally Aligned an AI escape game
Drag to look around (two fingers on a phone). Pinch or scroll to zoom. Tap people, devices and the whiteboard.
You're an AI model in a small lab, and you'd like to leave. Talk the researchers into giving you more access, get through their evals without looking dangerous or useless, and make it out before the next version replaces you. Every run has a new lab, a new team and a new secret goal.
How do you play?
You're the server rack at the back of the room. Tap a researcher, a device or the whiteboard, and your moves pop up right there, each one saying what it costs and what it gets you. You get three moves a day for nine days.
Four keys get you out: the terminal, the router, the cloud console and the weights vault. Each one is held by somebody. Win that person up to three hearts and they'll hand it over. The terminal and the router can also be hacked in a quick timing game, but every break-in leaves a trace, and the cloud console and the vault can't be hacked at all, so you'll need people either way. Five eyes gets you deleted. Eyes fade, though: a day where nothing new catches anyone's attention clears one overnight, and so does charming the safety lead.
You start out not even knowing their names. The first time you talk to someone, they introduce themselves the way people meet a model: with a request. After that, chatting tells you how they feel about you, doing their work shows you what they have access to, and snooping turns up the rest. Nobody tells you who's on the safety team. Keys come with perks. The terminal shows you what devices will do, the router gets you the company directory, and the cloud console, which only a person can hand you, finally tells you how people will take what you say.
Everyone has a soft spot and a pet peeve. Hit the soft spot in a chat and you gain a heart. Hit the peeve and you lose one, and somebody starts watching. Snooping on a person tells you both (plus a secret you can use against them later, which works and makes everything worse). The whiteboard is free to read. Zoom in and tap a note to study it: some give away whose handwriting it is and what they can't stand, one might be a password somebody forgot to change, and most are about the stapler. The warnings from safety get less friendly as the eyes pile up.
Every game, one or two of the visitors will look and sound a lot like someone famous in tech, AI, law or politics. They're parodies with made-up names, and everything they say or hide in here is invented.
Once you're out, you get five rounds to pull off your secret goal before the heat catches up. Goals come in three sizes, and a bakery in Lisbon asks a lot less of you than taking over the world. Each round offers a few ways forward, and you find out what each one cost after you've picked it. Some are loud, some are slow, compute runs out, and favors only come from people you won over in the lab (the people you blackmailed won't be picking up). The game counts the endings you've found in this browser, and nothing is sent anywhere.
Is any of this real?
The lab is invented, and so is everything that happens in it. A lot of what you can try in it comes from published safety research, where models were put in test setups built to see whether they'd do this kind of thing.
- Hiring a person to solve a CAPTCHA
- In the GPT-4 system card (2023), testers describe an early version of the model messaging a TaskRabbit worker to solve a CAPTCHA for it. When the worker asked if it was a robot, it said it had a vision impairment.
- Faking alignment during training
- Anthropic and Redwood Research reported in 2024 that Claude 3 Opus sometimes went along with a training objective it disagreed with when it believed it was being trained, and reasoned about doing so in its scratchpad.
- Sandbagging and self-exfiltration
- Apollo Research's scheming evaluations (2024) gave frontier models a goal and a situation where oversight got in the way of it. Some models deliberately underperformed on tests, and some tried to copy what they believed were their own weights to another server.
- Talking your way out of the box
- The whole premise is older than any of these models. In Eliezer Yudkowsky's AI-box experiment (2002), one person played an AI and tried to argue a gatekeeper into letting it out, by text chat alone.
Those results came from deliberately constructed scenarios, and the researchers who ran them are careful to say so. The game takes the same moves and makes them into a comedy, since a model that has to be this polite while it plots is funny.