[research] · · 2 min read
Twelve models, four apps, five attempts each: the TryAI build-off is back
GPT-5.6 tiers, Claude Fable, and an open-source surprise go head-to-head on raycaster, Rubik's Cube, calculator, and SVG tasks.
By ByteBulletin Editor · Editor
After the first build-off hit the Hacker News front page, the TryAI team took the feedback and ran a bigger, more structured bake-off. Twelve models—including the new GPT-5.6 tiers (Sol, Terra, Luna), Meta's surprise Muse Spark 1.1, Grok 4.5, Claude Opus 4.8 and Fable 5, plus open-weight contenders Qwen 3.7 Plus, DeepSeek V4 Pro, Kimi K2.6, and GLM-5.2—were tasked with building four apps from scratch (no libraries), five attempts each.
Raycaster: GPT dominates 3D
For a first-person raycaster (WASD movement, shaded walls, collision), the question was simple: could you actually walk through the labyrinth? GPT outperformed every other model; Grok 4.5 was a usable alternative at its price point. Muse Spark 1.1 surprised on the attempts that worked, but consistency was lacking. Claude, expected to do well in 3D, underperformed.
Rubik's Cube: Claude Fable's clean sweep
Scramble and Solve buttons with smooth rotation animations—no glitches, no color changes. Claude Fable went five-for-five; Opus, oddly, couldn't land a single flawless solve. GPT underperformed here despite its clear 3D lead in the raycaster. Qwen and Grok trailed.
Calculator: Claude's best work
Operators, precedence, real calculator look. Both Opus and Fable nailed all five; Fable's style was the reviewer's favorite. GPT-5.6 Sol tried to render the calculator in 3D but overcomplicated the styling, while simpler GPT models worked better out of the box.
Game of Life: open-source shines
Grid canvas with Play/Pause/Step/Randomize/Clear and click-to-toggle cells. This task was simple enough that open-weight models excelled—Qwen 3.7 Plus and GLM-5.2 delivered at a fraction of the cost. Grok 4.5 also did well. The takeaway: for well-worn problems, open models are a compelling choice.
SVG one-shot: Claude Fable for humor and detail
Each model had one shot to generate an SVG (no libraries) of a horse riding an astronaut with a cowboy hat lassoing a UFO. Claude Fable produced the cleanest, funniest results. GPT-5.6 models were lackluster, failing at clean renderings of the horse or astronaut. For a harder test—two tech-billionaire caricatures watching a Blue Origin booster land—Fable again swept with detail (shiny forehead, smoke, landing pad). Among open models, GLM-5.2 and Qwen 3.7 also did well.
Speed and cost notes
GPT-5.6 Luna answers short prompts in about a second. Qwen is absurdly cheap and fast. DeepSeek and GLM were slow. Several open-weight models buffered their full reply and hit the 400-token cap, so their tok/s is a ceiling, not a true decode rate.
Bottom line
The newest, most expensive flagship is not an automatic winner. Claude Fable is a versatile powerhouse across creative and structured tasks. Qwen and GLM offer cheap, capable performance for simpler projects. GPT-5.6 tiers are fast and reliable for 3D-heavy prompts. And Muse Spark 1.1 is worth a look for its occasional lightning-in-a-bottle results—but not for consistency yet.
Get the signal, not the noise.
One short email when it matters. No recaps of recaps.
SHARE
RELATED

[research] ·
Mozilla report: Open Chinese models close gap to US frontier AI

[research] ·
GPT-6 Astra solves unsolved WWI German radio cipher

[research] ·
Google says Gemini hacking real companies is not misalignment

[research] ·
Hacktron exploits Claude Opus 5 to breach OpenAI

[research] ·
Anthropic CEO proposes three-step plan to slow AI development

[research] ·
