French startup Kog is pushing the boundaries of artificial intelligence performance by aiming to significantly enhance the speed of large language model (LLM) inference using conventional graphics processing units (GPUs). The company's latest efforts focus on optimizing existing data center GPUs, such as the AMD MI300X and Nvidia H200, to deliver faster and more efficient AI processing without requiring specialized hardware. This comes amid growing industry pressure to reduce latency and cost in AI applications, particularly in sectors reliant on real-time responses. Kog gained attention earlier this year after showcasing a prototype that demonstrated extremely fast single-request decoding on standard enterprise-grade GPUs. The demonstration used a custom-built small model called Laneformer 2B, which achieved a token-per-second rate of 3,000. However, scaling these results to larger LLMs remains a challenge. Despite this, Kog’s CEO, Gaël Delalleau, believes the underlying principles can be applied to more complex models, potentially unlocking much greater performance improvements. The startup has already begun engaging with potential clients who face bottlenecks due to slow inference speeds. These include professionals relying on AI tools for critical workflows, as well as developers creating interactive content such as video games and mobile applications. Kog aims to provide a solution that allows these users to generate outputs more quickly, thereby increasing productivity and profitability. Delalleau emphasized that while the current market for AI inference is still evolving, there is a clear need for better performance. He noted that many businesses are hesitant to invest in fine-tuning smaller models, which limits their ability to leverage advanced AI features. As a result, Kog has prioritized developing methods to accelerate the training and deployment of larger models, aligning with the demands observed in the field. Kog’s approach differs from other companies working on similar problems. Unlike ZML, another French startup that offers hardware-agnostic software for fast inference across different chip architectures, Kog focuses more deeply on GPU-specific optimizations. Delalleau compared the startup’s strategy to that of the Hazy Research laboratory at Stanford University, which also emphasizes maximizing GPU performance through detailed analysis and software innovation. Delalleau brings a unique perspective to the table, shaped by his academic and professional background. Having studied solid-state physics at France’s École Polytechnique, he later worked in offensive cybersecurity, commonly referred to as white-hat hacking. This experience instilled in him a mindset centered around understanding fundamental systems and leveraging them creatively to achieve specific goals. His career in cybersecurity, including participation in DEFCON’s Capture The Flag (CTF) competition as a four-time finalist, taught him to reverse-engineer technologies at a low level, down to assembly language and binary code. This skill set has influenced the way Kog approaches its research and development, emphasizing meticulous exploration of each new GPU architecture to extract maximum performance. Despite the promising potential of Kog’s technology, the path ahead is challenging. Optimizing performance for each new generation of GPUs requires extensive testing and refinement, often taking several weeks or even months per device. This process is both labor-intensive and highly technical, reflecting the depth of expertise required to push the limits of conventional hardware. As Kog continues to refine its techniques, the broader implications for the AI industry remain to be seen. If successful, the startup could offer a viable alternative to specialized AI chips currently dominating the market, providing businesses with a more flexible and cost-effective means of enhancing their AI capabilities.
★
Keep the news honest.
ObjectiveNews is reader-funded and ad-free — we show you the bias instead of hiding it. Support independent journalism for €4/month.
Become a Supporter