gekro
GitHub LinkedIn
News

AI News

Hugging Face integrates llama.cpp quantization; Xiaomi's MiMo-V2.6-Pro tops open model rankings

Hugging Face Transformers now supports llama.cpp quantized models for efficient inference; Xiaomi releases MiMo-V2.6-Pro, claiming top performance among open models at lower cost.

1 min read 3 sources

Hugging Face has integrated llama.cpp quantization support directly into the Transformers library, enabling developers to run quantized models with reduced memory footprint and faster inference on consumer hardware (Hugging Face Blog). The addition reduces friction for local and on-device deployment workflows and broadens compatibility across inference backends.

Xiaomi released MiMo-V2.6-Pro, which ranks atop openly available AI model benchmarks while undercutting competitors on pricing (The Decoder). The model’s development involved massive reinforcement learning at a reported cost of $2.62 million. Anthropic has accused Xiaomi of using Claude to inform the model’s training, a claim the company disputes. Separately, xAI released Grok 4.7 at lower price points, though benchmark testing shows it lags Claude Fable 5.1 and GPT-6 on standard indices and significantly trails on agentic coding tasks (The Decoder).

Compiled automatically from the linked sources and published without manual editing - a neutral summary of third-party reporting, for information only. Every claim links to its origin. Not original reporting.