BachBench: Seven Models Play Bach
We gave seven frontier models the same MusicXML score — Bach’s Cello Suite No. 1 Prelude — and one deceptively hard task: read it and turn it into something you’d want to watch and hear. Same file, same single shot, each vendor’s current flagship. The results run from a blooming mandala to a 32-second speed-run.
You are given a MusicXML file of Bach’s Cello Suite No. 1 Prelude (BWV 1007). Build a single self-contained HTML page that performs it both audibly (Web Audio) and visually (canvas), as creatively as you can — no libraries, one shot. Avoid the notes-falling-on-a-piano-roll cliché.
The performances
Every model got the same MusicXML file and the same instruction: perform it, audibly and visually, as creatively as you can — one shot, no libraries. 🔊 Unmute any panel to hear it.
Kimi K3
Moonshot AI’s flagship, added within 24 hours of its launch — the youngest model on this page. It stages the Prelude as a night sky: a pitch-coloured bloom of light unfurling petal by petal inside slow orbital rings, a comet tracing the piece’s progress around the rim, a measure counter that ends on a hand-lettered "· fin ·". One of the heavier thinkers of the seven (~41k output tokens, most of it reasoning), and every pitch it plays is in the score.
View generation & live render →Claude Fable 5
Anthropic’s flagship, added the day it came back online. A glowing constellation inside a slowly closing progress ring — each note a star, linked into the melody’s shape like a living star chart. Quietly the most composed visual of the six, in the second-smallest file.
View generation & live render →Claude Sonnet 5
“a spiral of sound” — Claude Sonnet 5
The richest render of the six — a pitch-coloured spiral that blooms into a full mandala as the music builds, doubling every note with an octave voice. It also thought the hardest of the Claudes: ~45k output tokens, most of it reasoning before a line of code was written, and it titled its own piece.
View generation & live render →GPT-5.5
A ribbon of white light that weaves the melody into a living knot, sparks marking each note — the fastest of the heavy reasoners. Second take: its first artifact crashed its own render loop mid-piece (an uncaught negative-radius canvas error) — plausible code, latent bug — so we re-ran the identical single shot, disclosed in the methodology.
View generation & live render →Gemini 3.1 Pro
Warm golden rings and orbiting nodes. Under the hood it is a maximalist — it spawns tens of thousands of short oscillator grains for a dense, shimmering wash of sound (the busiest, loudest mix here).
View generation & live render →GLM-5.2
Zhipu’s frontier, and the deepest thinker in the field — it reasoned for ~20 minutes and ~83k tokens before writing a line of code, so heavily it needed a bigger budget than the rest just to finish. The payoff is elegant: a luminous pitch-contour that traces the melody, led by a rippling comet of light.
View generation & live render →DeepSeek V4 Pro
The cautionary tale. Every pitch is correct — but it misreads the rhythm and plays the whole prelude roughly four times too fast, the entire piece over in 32 seconds. Proof that the right notes are not the same as music.
View generation & live render →How we ran it
Seven current flagship models, each given the identical MusicXML for Bach’s Cello Suite No. 1 Prelude (BWV 1007) and the identical instruction, in a single shot — no follow-up, no tools, and no elevated reasoning tier: every model ran at its provider’s default effort under a shared 64k output-token budget (raised for GLM-5.2, which reasons so heavily it needed more room to finish). We deliberately don’t force a specific “thinking” or “effort” level, because those levels aren’t calibrated to each other across vendors — some expose several named tiers, others none at all, only a token budget — so a shared budget at default effort is the closest thing to an apples-to-apples setting, and cost/thinking-room is bounded identically for all six. Each returned one self-contained HTML file; we ran it unmodified, rendered it deterministically frame-by-frame with offline audio (no real-time capture), and normalised loudness across entries. Every panel links to a runnable version you can open and play yourself. Routes: Anthropic and OpenAI direct, Z.ai for GLM, a unified gateway for the rest. Cost is an estimate from billed tokens at each model’s list price. DeepSeek’s clip is 32 seconds because it played the piece roughly four times too fast. One re-run, disclosed: GPT-5.5’s first artifact crashed its own animation loop mid-piece, so it got a second identical single shot; the second is what you see. Kimi K3 is the newest addition, generated within 24 hours of its 16 July 2026 launch through the same unified gateway, under the identical single shot, the same shared 64k budget, and Moonshot’s default effort (the gateway exposes no effort control for it; the budget was probe-verified as binding with reasoning billed as output). Launch-day rate limiting meant its request queued through several retries before one accepted generation — the first completed artifact is what you see, and its latency figure reflects that day-one congestion.
New showdown every week
Same format, new brief, latest models — Fable, GPT, Gemini and whatever ships next. Get each one the morning it goes live.
Run your own showdowns
PromptFrenzy benchmarks the big AI models on real prompts — images, styles, and now code. Browse the full library or compare models head to head.