[launches] · · 2 min read
Kog aims to squeeze 30x faster LLM inference out of existing GPUs
The French startup is betting that deep software optimization can unlock far more performance from the datacenter GPUs enterprises already own.
By ByteBulletin Editors · Editorial Team
As the race for faster AI inference heats up, French startup Kog is taking a different path: instead of building custom silicon like Cerebras, it's focused on squeezing more performance out of the GPUs that enterprises already own. The company's promise — up to 30x faster LLM inference — has already captured attention, drawing 200 tangible business leads after a tech preview that hit the front page of Hacker News in May.
Kog's early demo showed an impressive 3,000 tokens per second for a single request, but that was achieved with a small, purpose-built 2B-parameter model, Laneformer 2B, which is now open sourced. The real challenge lies in scaling this performance to large language models, which are far more demanding on memory bandwidth and compute. CEO Gaël Delalleau, however, is confident that the same approach will work, arguing that newer GPUs have plenty of memory bandwidth that software optimization can unlock.
The company's approach is deeply hands-on, rooted in Delalleau's background in solid-state physics and offensive cybersecurity. He describes a mindset of understanding the laws of the GPU at a low level, even down to assembly language, in order to use it for purposes it wasn't necessarily designed for. This means dedicating weeks or months to each new GPU model, which limits how many chips Kog can support with its team of 11. In the long run, the company hopes to automate this process with agent-based pipelines.
Initially, Kog plans to target software engineering workflows, where users of tools like Claude Code often face long waits for results. The company has also identified design partners building prompt-to-game and prompt-to-app platforms, where faster inference directly translates to more revenue. Delalleau says Kog has learned that its customers aren't interested in fine-tuning small models, so the company is now focused on accelerating larger models to meet that demand.
The startup has a clear roadmap: implement a major model at 10x speed by September, demonstrate customer traction, and then raise a Series A. With backing from French institutions like Bpifrance and support from Scaleway, Kog also benefits from Europe's push for technological sovereignty. The next few months will be critical as Kog attempts to prove that its approach works on the large models that matter in production.
SHARE
RELATED

[launches] ·
Apple Trains Custom AI Model with Alibaba for China Launch
A rare US-China partnership could make Apple the first American company with a government-approved proprietary AI model in China.

[launches] ·
OpenAI replaces CRO after nine months as executive shake-up intensifies
Dali Rajic, former Wiz president, takes over as OpenAI's chief revenue officer amid a broader leadership reshuffle.

[launches] ·
Twitch’s AI Training Opt-Out: Amazon’s Default-Yes Policy Sparks Backlash
Twitch now lets streamers opt out of Amazon’s generative AI training, but the default-opt-in approach has ignited a community firestorm.