[models] · · 1 min read
Writer launches Palmyra X6, a post-trained model built to slash token costs
The new flagship model, based on Z.ai's open-source GLM-5.2, pairs with an upgraded harness to cut deployment costs by up to 50% for enterprise customers.
By ByteBulletin Editors · Editorial Team
Enterprises are increasingly feeling the sting of AI deployment costs, and Writer is betting that a combination of a new model and a more efficient "harness" can ease the pain. On Thursday, the company launched Palmyra X6, a flagship model built as a post-training variation on Z.ai's open-source GLM-5.2. The model is designed to deliver deployment-ready capabilities at a significantly lower price point, and Writer estimates that, together with changes to its agentic harness, customers could see costs drop by as much as 50% for basic tasks.
The launch comes amid growing customer frustration with the rising cost of AI deployments, a sentiment Writer CEO May Habib says is driving a shift in enterprise priorities. "The enterprise is absolutely sick of chasing the next benchmark," Habib told TechCrunch. "They want flattening cost, and it seems like nobody can deliver that."
Writer's strategy hinges on optimizing the harness—the software layer that orchestrates model calls and manages agent workflows—rather than relying solely on model selection. A recent paper from Writer researchers found that small changes in harness efficiency were often a more reliable way to reduce costs than switching models, with costs falling an average of 40% across their testing. The harness, they argue, is "the one component whose efficiency multiplies across every model an organization runs—present and future."
The upgraded harness emphasizes complex, multi-step tasks, executing them faster and with fewer tokens. For customers, the experience remains model-agnostic: Palmyra X6 sits alongside other Writer models or external models imported via Azure or Amazon Bedrock. Habib also sees the cost-cutting push as fueling broader distrust toward major AI labs, which she claims have a financial incentive to drive up token usage. "The cost explosion here is just unprecedented for customers, and so is the degree to which CIOs are giving up on the labs," she said.
Palmyra X6 and the updated harness are available to Writer clients starting Thursday.
SHARE
RELATED

[models] ·
Meta doubles down on open-weight AI with Muse Glimmer release and Zuckerberg manifesto
Meta releases a 30B-parameter open model for local use, promises Spark 1.2 weights, and makes a philosophical case for decentralized AI.

[models] ·
Meta's Glimmer model brings 'personal superintelligence' to your laptop
With the open-weight Muse Glimmer, Meta offers a local, privacy-conscious AI agent for consumer hardware—and a glimpse of where it draws the line between open and closed AI.

[models] ·
Escha Labs ships 2-bit Qwen3.6-35B-A3B build that runs on a 24GB GPU
A 12.3GB MoE quant packs a 35B model onto consumer cards, with quality within noise of FP8 on most benchmarks.