French startup Kog targets faster LLM inference on existing GPUs with low-level optimization
Kog claims up to 30x speedups on standard datacenter GPUs, but has only demonstrated the approach on a small open-source model so far.
1 source · cross-referenced
- Kog, a French startup, claims its software can deliver up to 30x faster LLM inference on standard GPUs like AMD MI300X and Nvidia H200 by optimizing at the hardware level.
- The company has only demonstrated the approach on an open-source 2B-parameter model (Laneformer 2B), achieving 3,000 tokens per second per request.
- Kog’s CEO argues GPUs remain well-suited for inference and attributes skepticism to a misconception about their decoding capabilities.
- The startup is now focused on scaling its method to larger models and expects to demonstrate a 10x speed improvement by September 2026.
Kog, a French startup, argues that conventional GPUs remain underutilized for AI inference and is betting on low-level software optimization to unlock more performance from existing hardware. The company’s approach targets the inference bottleneck that has constrained agentic and professional AI workflows, according to CEO Gaël Delalleau.
Kog’s initial demonstration used the open-source Laneformer 2B model on AMD MI300X and Nvidia H200 GPUs, achieving 3,000 tokens per second per request. The company claims this validates its broader promise of “30x faster LLM inference,” though it has not yet demonstrated the same speedups on larger models.
Delalleau contends that skepticism about GPUs’ suitability for decoding is a misconception, pointing to increasing memory bandwidth in newer GPUs as an untapped resource. He describes Kog’s method as akin to reverse-engineering hardware behavior down to assembly and binary code to maximize efficiency.
The startup’s engineering process is labor-intensive: for each new GPU, Kog dedicates weeks or months to deep hardware research before deploying its optimizations. With a team of 11, this limits the number of supported chips in the near term.
Kog has generated 200 business leads and expects software engineers—particularly users frustrated by slow tool-assisted workflows—to be early adopters. The company also works with design partners in gaming and app generation, where faster inference translates directly to revenue.
Kog’s seed round was co-led by Varsity VC, with support from Scaleway and France’s Bpifrance and French Tech 2030 program. The startup aims to demonstrate a 10x speed improvement on a major model by September 2026 to secure Series A funding and prove broader applicability.
Kog is not alone in pursuing software-driven GPU optimization; competitors like ZML offer hardware-agnostic software that bypasses Nvidia’s CUDA. However, Delalleau positions Kog’s work as deeper and more focused, drawing a parallel to Stanford’s Hazy Research lab.
- Aug 14, 2026 · Hugging Face
Hugging Face finds Chinese labs lead open model releases at frontier scale while U.S. hardware vendors dominate new model uploads
Trust79 - Aug 14, 2026 · TechCrunch — AI
Google adds toggle to remove visible AI watermarks from images, video, and audio
Trust79 - Aug 14, 2026 · TechCrunch — AI
Writer unveils Palmyra X6, a post-trained variant of Z.ai’s GLM-5.2, and upgrades its agentic harness to cut token costs
Trust78