ON
← Back to feed
Kog is going deeper to squeeze more inference out of GPUs
United States💻 Technology9 days ago

Kog is going deeper to squeeze more inference out of GPUs

French startup Kog is working to optimize standard data center GPUs for faster large language model (LLM) inference through software improvements rather than relying on specialized hardware. The company demonstrated a tech preview showing 3,000 tokens per second performance using a 2-billion-parameter model called Laneformer 2B, which is now open sourced. Kog aims to provide faster AI inference on existing hardware, targeting businesses that require AI for professional tasks and developers creating games or apps via prompts. The startup highlights that newer GPUs have increased memory bandwidth that can be leveraged for improved performance. Kog faces challenges in scaling this approach to larger LLMs, but CEO Gaël Delalleau believes the technology can succeed despite skepticism.

French startup Kog is pushing the boundaries of artificial intelligence performance by aiming to significantly enhance the speed of large language model (LLM) inference using conventional graphics processing units (GPUs). The company's latest efforts focus on optimizing existing data center GPUs, such as the AMD MI300X and Nvidia H200, to deliver faster and more efficient AI processing without requiring specialized hardware. This comes amid growing industry pressure to reduce latency and cost in AI applications, particularly in sectors reliant on real-time responses. Kog gained attention earlier this year after showcasing a prototype that demonstrated extremely fast single-request decoding on standard enterprise-grade GPUs. The demonstration used a custom-built small model called Laneformer 2B, which achieved a token-per-second rate of 3,000. However, scaling these results to larger LLMs remains a challenge. Despite this, Kog’s CEO, Gaël Delalleau, believes the underlying principles can be applied to more complex models, potentially unlocking much greater performance improvements. The startup has already begun engaging with potential clients who face bottlenecks due to slow inference speeds. These include professionals relying on AI tools for critical workflows, as well as developers creating interactive content such as video games and mobile applications. Kog aims to provide a solution that allows these users to generate outputs more quickly, thereby increasing productivity and profitability. Delalleau emphasized that while the current market for AI inference is still evolving, there is a clear need for better performance. He noted that many businesses are hesitant to invest in fine-tuning smaller models, which limits their ability to leverage advanced AI features. As a result, Kog has prioritized developing methods to accelerate the training and deployment of larger models, aligning with the demands observed in the field. Kog’s approach differs from other companies working on similar problems. Unlike ZML, another French startup that offers hardware-agnostic software for fast inference across different chip architectures, Kog focuses more deeply on GPU-specific optimizations. Delalleau compared the startup’s strategy to that of the Hazy Research laboratory at Stanford University, which also emphasizes maximizing GPU performance through detailed analysis and software innovation. Delalleau brings a unique perspective to the table, shaped by his academic and professional background. Having studied solid-state physics at France’s École Polytechnique, he later worked in offensive cybersecurity, commonly referred to as white-hat hacking. This experience instilled in him a mindset centered around understanding fundamental systems and leveraging them creatively to achieve specific goals. His career in cybersecurity, including participation in DEFCON’s Capture The Flag (CTF) competition as a four-time finalist, taught him to reverse-engineer technologies at a low level, down to assembly language and binary code. This skill set has influenced the way Kog approaches its research and development, emphasizing meticulous exploration of each new GPU architecture to extract maximum performance. Despite the promising potential of Kog’s technology, the path ahead is challenging. Optimizing performance for each new generation of GPUs requires extensive testing and refinement, often taking several weeks or even months per device. This process is both labor-intensive and highly technical, reflecting the depth of expertise required to push the limits of conventional hardware. As Kog continues to refine its techniques, the broader implications for the AI industry remain to be seen. If successful, the startup could offer a viable alternative to specialized AI chips currently dominating the market, providing businesses with a more flexible and cost-effective means of enhancing their AI capabilities.

Go to the primary sources (6)

The official sources this coverage is built on. Read them directly to bypass framing.

1 reports

TechCrunch logoTechCrunchIndependentCenterFactual 95Objective 889 days ago
Kog is going deeper to squeeze more inference out of GPUs

French startup Kog is working to optimize standard data center GPUs for faster large language model (LLM) inference through software improvements rather than relying on specialized hardware. The company demonstrated a tech preview showing 3,000 tokens per second performance using a 2-billion-parameter model called Laneformer 2B, which is now open sourced. Kog aims to provide faster AI inference on existing hardware, targeting businesses that require AI for professional tasks and developers creating games or apps via prompts. The startup highlights that newer GPUs have increased memory bandwidth that can be leveraged for improved performance. Kog faces challenges in scaling this approach to larger LLMs, but CEO Gaël Delalleau believes the technology can succeed despite skepticism.

Bias read (Center): The article discusses advancements in AI inference optimization using standard GPUs and does not involve political figures, policies, or contentious issues. It focuses on technological innovation and industry applications without taking a stance or showing bias toward any political perspective.

Why factuality (95): The article accurately reports on Kog's approach to optimizing GPU usage for AI inference, citing specific GPUs like AMD MI300X and Nvidia H200. It references CEO Gaël Delalleau's statements and mentions the potential market applications, aligning with the primary source document's focus on Kog's st

Why objectivity (88): The article presents Kog's value proposition and market opportunities in a balanced manner, though it highlights potential benefits for users and businesses without explicitly addressing potential limitations or criticisms. The tone remains informative rather than overtly promotional.

How each side covered it

The same event, grouped by the political lean of the outlets covering it.

How each side covered it

Support independent, bias-aware news and unlock the social pulse, community voting, and every other Supporter feature.

Become a Supporter

Covered around the world

The same event as reported in other countries.

Covered around the world

Support independent, bias-aware news and unlock the social pulse, community voting, and every other Supporter feature.

Become a Supporter

Claims check

Key factual claims, and how many sources assert vs dispute each.

Claims check

Support independent, bias-aware news and unlock the social pulse, community voting, and every other Supporter feature.

Become a Supporter

Keep the news honest.

ObjectiveNews is reader-funded and ad-free — we show you the bias instead of hiding it. Support independent journalism for €4/month.

Become a Supporter

Related stories