gekro
GitHub LinkedIn
News

AI News

OpenAI's Jalapeño inference chip outperforms Nvidia on throughput and efficiency; IBM releases Granite 4.2 open-weight models

OpenAI's custom Jalapeño chip demonstrated superior inference throughput and energy efficiency versus Nvidia's latest offerings, while IBM published technical details on Granite 4.2 LLMs.

1 min read 5 sources

OpenAI presented benchmark results for Jalapeño, its first in-house inference chip, at the Hot Chips conference. Testing by SemiAnalysis on the InferenceX benchmark showed Jalapeño achieved higher tokens per user and throughput per kilowatt than competing inference accelerators (The Decoder). The chip registered both more tokens per second and improved energy efficiency compared to currently available state-of-the-art solutions (TechCrunch), positioning it as a purpose-built system for inference scaling rather than training workloads.

IBM published a technical breakdown of how its Granite 4.2 LLM series was constructed (Hugging Face Blog). Separately, Hugging Face documented a quantization technique called Quantization-Aware Healing that produced a compressed 4-bit model exceeding the performance of its full-precision baseline (Hugging Face Blog). Google launched Gemini Enterprise for Legal, which integrates with contract and legal systems through MCP connectors and enables AI agents for document review and research tasks (The Decoder).

Compiled automatically from the linked sources and published without manual editing - a neutral summary of third-party reporting, for information only. Every claim links to its origin. Not original reporting.