Kog Enables Processing of 3,000 Tokens Per Second with Standard GPUs
French startup Kog has announced a strategy to enhance AI inference speeds using standard data center GPUs. Kog has adopted an approach that involves designing software and model architecture together to maximize the utilization of NVIDIA and AMD GPUs. The company stated that it has integrated runtime, GPU kernels, and model architecture into an optimized pipeline to improve response speeds to requests from AI agents. According to a technology preview of the Kog inference engine released on May 28, it can process 3,000 output tokens per second per request using eight AMD MI300X GPUs, while eight NVIDIA H200 GPUs under the same conditions recorded 2,100 output tokens per second. Kog is primarily targeting AI coding agents and agent-based workflows. The Laneformer 2B model, released on Hugging Face on June 24, has 2.3 billion parameters and demonstrated performance of 45.1% on HumanEval+ and 51.6% on MBPP+. Kog reported in an August LinkedIn post that it achieved 2,857 tokens per second in a live demo. However, there is controversy regarding the interpretation of speed figures, and it has been pointed out that comparisons are difficult due to variations in model size and hardware. Kog's goal is to achieve low latency on existing GPU infrastructure.
-- Price
This content is provided for general informational purposes only and doesn't constitute financial, investment, legal, or tax advice. Any events, rewards, online promotions, or related information mentioned herein should not be considered a recommendation, solicitation, or invitation to purchase, sell, trade, or otherwise deal in any crypto assets. Crypto assets are highly volatile and may result in loss. The availability of WEEX services, products, and related events may vary by region. You are responsible for ensuring that your participation is in accordance with applicable local laws and regulations.
You may also like

Tether Blocks 42 Million USDT Before Court Ruling

$317 Billion Stablecoins Become a New Demand Layer for Short-Term Treasuries

Last Night, Silicon Valley Experienced a Battle of AI Titans

Solana inflation cut is premature, SOL Strategies CEO says

Norway Arrests Russian Vessel at Naftogaz's Request Over $4.22 Billion Debt

What the $344M crypto political spending spree wants from Congress next

OpenAI Announces Astra as the First Model with Critical Cybersecurity Capabilities

SEC novel ETF review draws opposition from crypto firms

AI: The Next Billion Crypto Users Will Not Be Human

Prospect Markets, Crypto.com seal deal for U.S. prediction markets platform

Ukraine and Lithuania Strengthen Cooperation to Save Grain Exports

Google Pics Brings AI Images to Workspace, Challenging Canva and Adobe

Flop Labs Launches tclk/1 Protocol for Trustless AI Agent Transactions

Ukraine Needs $27 Billion for Defense

Thai businessmen sue Tether over 42.4M USDT freeze

How recovery of 61 BTC unlocked a potential $432M treasure hunt for early Bitcoin users

Apple Requests Court to Halt OpenAI Hardware Project

Predict Developer Dashboard Launched, Supports Application Creation and Rate Limit Management

Silhouette Launches RFQ System to Support On-Chain Trading of xStocks

Attacks on Logistics Infrastructure Change Job Market Geography: Where Demand for Workers is Growing

Japanese Ministries Request Record Budget as National Bond Yields Rise to 3%

Financial Services Agency Requests Simplification of Stablecoin Taxation|Facilitating Payments Over 1 Million Yen

DAXA Requests Investigation into Four Unreported Overseas Cryptocurrency Exchanges

Polymarket Discusses $1 Billion Funding, Valuation of 29 Trillion Won Mentioned

Peach Limits P2P Bitcoin Sales Under New Rules

Sony Claims Digital Game Ownership Is Misleading in Court

Hacking: Belgium demands crypto addresses from illegal sites

ARCA Expands the Obligation to Issue Electronic Receipts

HYPE faces $36M team withdrawal on September 6










