[research] · · 2 min read
Twelve models, four apps, five attempts each: the TryAI build-off is back
GPT-5.6 tiers, Claude Fable, and an open-source surprise go head-to-head on raycaster, Rubik's Cube, calculator, and SVG tasks.
By ByteBulletin Editors · Editorial Team
After the first build-off hit the Hacker News front page, the TryAI team took the feedback and ran a bigger, more structured bake-off. Twelve models—including the new GPT-5.6 tiers (Sol, Terra, Luna), Meta's surprise Muse Spark 1.1, Grok 4.5, Claude Opus 4.8 and Fable 5, plus open-weight contenders Qwen 3.7 Plus, DeepSeek V4 Pro, Kimi K2.6, and GLM-5.2—were tasked with building four apps from scratch (no libraries), five attempts each.
Raycaster: GPT dominates 3D
For a first-person raycaster (WASD movement, shaded walls, collision), the question was simple: could you actually walk through the labyrinth? GPT outperformed every other model; Grok 4.5 was a usable alternative at its price point. Muse Spark 1.1 surprised on the attempts that worked, but consistency was lacking. Claude, expected to do well in 3D, underperformed.
Rubik's Cube: Claude Fable's clean sweep
Scramble and Solve buttons with smooth rotation animations—no glitches, no color changes. Claude Fable went five-for-five; Opus, oddly, couldn't land a single flawless solve. GPT underperformed here despite its clear 3D lead in the raycaster. Qwen and Grok trailed.
Calculator: Claude's best work
Operators, precedence, real calculator look. Both Opus and Fable nailed all five; Fable's style was the reviewer's favorite. GPT-5.6 Sol tried to render the calculator in 3D but overcomplicated the styling, while simpler GPT models worked better out of the box.
Game of Life: open-source shines
Grid canvas with Play/Pause/Step/Randomize/Clear and click-to-toggle cells. This task was simple enough that open-weight models excelled—Qwen 3.7 Plus and GLM-5.2 delivered at a fraction of the cost. Grok 4.5 also did well. The takeaway: for well-worn problems, open models are a compelling choice.
SVG one-shot: Claude Fable for humor and detail
Each model had one shot to generate an SVG (no libraries) of a horse riding an astronaut with a cowboy hat lassoing a UFO. Claude Fable produced the cleanest, funniest results. GPT-5.6 models were lackluster, failing at clean renderings of the horse or astronaut. For a harder test—two tech-billionaire caricatures watching a Blue Origin booster land—Fable again swept with detail (shiny forehead, smoke, landing pad). Among open models, GLM-5.2 and Qwen 3.7 also did well.
Speed and cost notes
GPT-5.6 Luna answers short prompts in about a second. Qwen is absurdly cheap and fast. DeepSeek and GLM were slow. Several open-weight models buffered their full reply and hit the 400-token cap, so their tok/s is a ceiling, not a true decode rate.
Bottom line
The newest, most expensive flagship is not an automatic winner. Claude Fable is a versatile powerhouse across creative and structured tasks. Qwen and GLM offer cheap, capable performance for simpler projects. GPT-5.6 tiers are fast and reliable for 3D-heavy prompts. And Muse Spark 1.1 is worth a look for its occasional lightning-in-a-bottle results—but not for consistency yet.
SHARE
RELATED

[research] ·
Google Warns of 'Vishing' Attacks Targeting Financial Firms with Extortion Demands
Hackers are using phone calls to trick employees at major investment firms into handing over credentials, then extorting them for millions.

[research] ·
New Research Predicts LLM Inference Latency at the Edge, Aiming for Smarter Offloading
A new arXiv paper proposes a method to forecast LLM inference latency before deployment, which could make edge-device offloading decisions far more reliable.

[research] ·
Google’s AI Leadership Shake-Up: Turmoil or a Strategic Pivot?
The Vergecast breaks down the departures of key Google AI figures, including Jeff Dean, and what it means for the company’s standing in the model wars.
