Overpromise an AI lab tycoon
Built by Claude Opus 5.5
You run a small AI lab. Investors fund what people believe your model can do, and people believe what you tell them on stage. Overpromise is about the distance between those two numbers, and how long you can keep it open before somebody checks.
How do you play?
The campus is the scoreboard. Solid floors on your tower are the model's real score, and the pink inflatable floors stacked on the roof are the hype. Rival labs stand on the skyline at the same scale, so you can see who's actually taller (it's rarely you, at first).
Hire researchers and sign for GPUs, and the next model trains by itself. When it's ready, hold a keynote (you get one every 40 days). Pick a promise, modest, bold or wild, then run the live demo: a cursor sweeps across the screen and you tap while the model's answer is right. A clean demo inflates more floors. A bigger promise inflates a lot more, shrinks the target on stage, and sets a deadline you'll have to meet.
More hype means a higher valuation, and a higher valuation means a bigger funding round. People do check, though. A reporter asks for a week with the model, a bank runs its pilot on real work, or a test turns up that nobody could have trained on. If your claims have drifted too far past the real score, the bubble pops. The pink floors burst, trust drops and users leave. The game never tells you how far is too far. The reporters gathering outside your lobby are a hint.
Shipping without red-teaming is faster and cheaper, and it tends to come back as an incident a few weeks later. Incidents bring regulators. If trust hits zero the board fires you, and a month without cash ends the company. You win by taking the lab public. After your last private round the board gives you about 240 days to file, and the IPO auditors test the product against every claim in the filing, so the gap has to be closed by then. Shipping quietly and walking your claims back on a podcast both help. Every lab you found earns credibility for perks in the next one. The rivals get a bit tougher each time you ring the bell.
It saves in this browser and sends nothing anywhere.
Is any of this real?
The lab, the rivals and the people are invented. A few visitors look a lot like famous tech figures. They're parodies with made-up names, and I kept the jokes to their public personas. The pattern underneath is much older than AI.
- The hype cycle
- Gartner analyst Jackie Fenn introduced the hype cycle in a 1995 research note. Two of its stages are called the Peak of Inflated Expectations and the Trough of Disillusionment, and a bad run in this game follows that shape pretty closely.
- Held-out benchmarks
- In 2024, researchers at Scale AI wrote GSM1k, a fresh set of grade school math problems in the style of the popular GSM8k test. Some model families scored up to 13% worse on the new set, which looked like partial memorization of the old one (the paper calls it systematic overfitting), while frontier models showed few signs of it. That study is where the "train on the test" button comes from, and why held-out tests punish it.