OpenAI presented benchmark results for Jalapeño, its first in-house inference chip, at the Hot Chips conference. Testing by SemiAnalysis on the InferenceX benchmark showed Jalapeño achieved higher tokens per user and throughput per kilowatt than competing inference accelerators (The Decoder). The chip registered both more tokens per second and improved energy efficiency compared to currently available state-of-the-art solutions (TechCrunch), positioning it as a purpose-built system for inference scaling rather than training workloads.
IBM published a technical breakdown of how its Granite 4.2 LLM series was constructed (Hugging Face Blog). Separately, Hugging Face documented a quantization technique called Quantization-Aware Healing that produced a compressed 4-bit model exceeding the performance of its full-precision baseline (Hugging Face Blog). Google launched Gemini Enterprise for Legal, which integrates with contract and legal systems through MCP connectors and enables AI agents for document review and research tasks (The Decoder).