[models] · · 1 min read
ByteDance Is Training a 10 Trillion-Parameter Model, Leaning Into Scale to Chase Anthropic
The TikTok parent is reportedly pre-training a massive model that could rival Anthropic's Mythos 5, signaling a new phase in the global AI race.
By ByteBulletin Editors · Editorial Team
ByteDance is reportedly in the early stages of training a massive AI model with up to 10 trillion parameters—three times larger than the biggest Chinese model released to date—according to three people with knowledge of the matter. The move positions the TikTok parent as the most ambitious Chinese lab yet in the race to match or exceed the top US models from Anthropic.
The model, being trained by ByteDance's Seed team, is currently in pre-training, a phase that typically lasts three to six months, followed by fine-tuning and potential release. The exact parameter count could change, but the reported scale signals a bet on sheer size as a path to frontier capability.
Industry estimates suggest Anthropic's Mythos 5 has around 8 trillion parameters, and its Fable 5 around 5 trillion. While parameter count sets fundamental capacity limits, capability also depends on data quality and training methods—factors that US labs have long used to maintain their edge.
SOURCES
SHARE
RELATED

[models] ·
Escha Labs ships 2-bit Qwen3.6-35B-A3B build that runs on a 24GB GPU
A 12.3GB MoE quant packs a 35B model onto consumer cards, with quality within noise of FP8 on most benchmarks.

[models] ·
Alibaba’s Qwen3.8-Max goes open-weight, claiming Claude-rivaling performance
Alibaba releases its largest open-weight model yet, Qwen3.8-Max, claiming it rivals Anthropic’s Claude Fable 5, with weights due next week.

[models] ·
OpenAI slashes GPT-5.6 Luna and Terra prices up to 80%, launches faster Sol tier
New API pricing makes Luna 80% cheaper and Terra 20% cheaper, while Sol's Fast mode delivers 2.5x speed at double the price.
