gekro
GitHub LinkedIn
News

AI News

llama.cpp v0.6.0 ships 3x Metal speedup; open-weight decision models target agent routing

llama.cpp v0.6.0 delivers 3x Apple Metal throughput with a new batch API; Cloudflare and Amazon ship open-weight decision models for agent routing.

2 min read 5 sources

llama.cpp released v0.6.0 on October 5, 2026 (llama.cpp GitHub), with the headline change being new Apple Metal MMA kernels delivering approximately 3x faster matrix multiplication on Apple GPUs for speculative and batched decoding (AI Weekly). The release introduces the llama_batch_ext extended-batch API, replacing the older llama_batch interface to support mixed token and embedding inputs alongside multi-token prediction draft-model state (llama.cpp GitHub). Day-one model support includes GLM-5.3-Flash, a 320B text-and-vision hybrid, and the Clef decision model via the vision-plus-text path; MTP speculative decoding for Qwen4Exp yields approximately a 1.5x decode speedup on DGX Spark hardware (AI Weekly). The Metal throughput gain has direct implications for practitioners running inference on Apple Silicon, where single-user throughput is otherwise constrained by unified-memory bandwidth.

vLLM published v0.31.0 on October 5, incorporating 717 commits from 307 contributors, 96 of them new to the project (vLLM GitHub). The release defaults to SM100-optimized FlashMLA attention with NVFP4 compressed KV cache for DeepSeek-V4.1-Flash on NVIDIA Blackwell hardware, and enables DeepGEMM sparse MQA logits for the V4.1 architecture by default (vLLM GitHub). These changes target practitioners serving the 552B DeepSeek-V4.1-Flash at production throughput on next-generation NVIDIA GPUs, bringing 4-bit quantization to what the project describes as a production-ready baseline for Blackwell deployments.

On October 1, 2026, Cloudflare and Amazon independently released open-weight models designed to return typed decisions rather than free text, a format intended for the routing and classification layer of agent pipelines. Cloudflare released Clef, a 27B-parameter model built on Qwen3.8-27B, and the smaller Clef-Flash at 9B; both carry an Apache 2.0 license with weights on Hugging Face and run on Workers AI, with Clef-Flash reaching a median decision latency of 38.8 milliseconds (MarkTechPost). Amazon Strands Labs published Strands Decider 2B on the same day, a 2B-parameter model based on Qwen3.5-2B with an added pointer head that selects from predefined options rather than generating text, reaching approximately 106 milliseconds on consumer hardware, also under Apache 2.0 with weights and training scripts on Hugging Face (AI Tech Daily). The simultaneous releases from two major cloud providers suggest decision models are emerging as a distinct infrastructure category for agent system designers who need low-latency, deterministic routing without the overhead of full text-generation inference.

Compiled automatically from the linked sources and published without manual editing - a neutral summary of third-party reporting, for information only. Every claim links to its origin. Not original reporting.