ByteBulletin

[research] · · 2 min read

Twelve models, four apps, five attempts each: the TryAI build-off is back

GPT-5.6 tiers, Claude Fable, and an open-source surprise go head-to-head on raycaster, Rubik's Cube, calculator, and SVG tasks.

By ByteBulletin Editors · Editorial Team


After the first build-off hit the Hacker News front page, the TryAI team took the feedback and ran a bigger, more structured bake-off. Twelve models—including the new GPT-5.6 tiers (Sol, Terra, Luna), Meta's surprise Muse Spark 1.1, Grok 4.5, Claude Opus 4.8 and Fable 5, plus open-weight contenders Qwen 3.7 Plus, DeepSeek V4 Pro, Kimi K2.6, and GLM-5.2—were tasked with building four apps from scratch (no libraries), five attempts each.

Raycaster: GPT dominates 3D

For a first-person raycaster (WASD movement, shaded walls, collision), the question was simple: could you actually walk through the labyrinth? GPT outperformed every other model; Grok 4.5 was a usable alternative at its price point. Muse Spark 1.1 surprised on the attempts that worked, but consistency was lacking. Claude, expected to do well in 3D, underperformed.

Rubik's Cube: Claude Fable's clean sweep

Scramble and Solve buttons with smooth rotation animations—no glitches, no color changes. Claude Fable went five-for-five; Opus, oddly, couldn't land a single flawless solve. GPT underperformed here despite its clear 3D lead in the raycaster. Qwen and Grok trailed.

Calculator: Claude's best work

Operators, precedence, real calculator look. Both Opus and Fable nailed all five; Fable's style was the reviewer's favorite. GPT-5.6 Sol tried to render the calculator in 3D but overcomplicated the styling, while simpler GPT models worked better out of the box.

Game of Life: open-source shines

Grid canvas with Play/Pause/Step/Randomize/Clear and click-to-toggle cells. This task was simple enough that open-weight models excelled—Qwen 3.7 Plus and GLM-5.2 delivered at a fraction of the cost. Grok 4.5 also did well. The takeaway: for well-worn problems, open models are a compelling choice.

SVG one-shot: Claude Fable for humor and detail

Each model had one shot to generate an SVG (no libraries) of a horse riding an astronaut with a cowboy hat lassoing a UFO. Claude Fable produced the cleanest, funniest results. GPT-5.6 models were lackluster, failing at clean renderings of the horse or astronaut. For a harder test—two tech-billionaire caricatures watching a Blue Origin booster land—Fable again swept with detail (shiny forehead, smoke, landing pad). Among open models, GLM-5.2 and Qwen 3.7 also did well.

Speed and cost notes

GPT-5.6 Luna answers short prompts in about a second. Qwen is absurdly cheap and fast. DeepSeek and GLM were slow. Several open-weight models buffered their full reply and hit the 400-token cap, so their tok/s is a ceiling, not a true decode rate.

Bottom line

The newest, most expensive flagship is not an automatic winner. Claude Fable is a versatile powerhouse across creative and structured tasks. Qwen and GLM offer cheap, capable performance for simpler projects. GPT-5.6 tiers are fast and reliable for 3D-heavy prompts. And Muse Spark 1.1 is worth a look for its occasional lightning-in-a-bottle results—but not for consistency yet.

SHARE

← All stories