AI Model Showdowns

One prompt, every frontier model, zero edits. We publish exactly what each model returns — the brilliant, the broken, and the gloriously weird.

Model Arena: Misaligned

Model Arena: Misaligned

7 frontier models, 36 games of hidden-saboteur deduction. The leaderboard: Claude Opus 4.8 is the best liar (88%); Gemini 3.1 Pro is the best lie-detector (83%).

BachBench: Seven Models Play Bach

BachBench: Seven Models Play Bach

We gave seven frontier models the same MusicXML score — Bach’s Cello Suite No. 1 Prelude — and one deceptively hard task: read it and turn it into something you’d want to watch and hear. Same file, same single shot, each vendor’s current flagship. The results run from a blooming mandala to a 32-second speed-run.

Build Your Dream Home

Build Your Dream Home

We gave eight AI models the same 21 materials, the same 48-cube grid, and one brief: build the home YOU would most want to live in. Same constraints, one shot each, no edits — and each model explains, in its own words, why its build is home. The choices say as much about the models as the builds do.

Cyberpunk Alley in Three.js

Cyberpunk Alley in Three.js

We gave eight AI models one prompt and one shot: write a complete Three.js scene from scratch — a neon-lit cyberpunk alley at night in the rain — with a locked-down 10-second camera dolly so every render lines up frame for frame. No tools, no retries, no edits. Every scene below is the model’s own code running in real time, from a polished neon corridor to one that never compiles. Which one nails the brief? Your call.

Pelican on a Bicycle

Pelican on a Bicycle

We asked the top frontier AI models — launch-day Claude Fable 5, GPT-5.5 Pro and Gemini 3.1 Pro — to draw a pelican riding a bicycle as SVG code. Same prompt, one shot, no edits. Then we made them animate it.