Aleph Alpha has released Kolibri, a 78.1B-parameter English-German Mixture-of-Experts model that activates only 3.46B parameters per token (MarkTechPost). The model supports a 1M-token context window and per-request reasoning effort, with Apache 2.0 licensed FP8 weights capable of running on a single B200 or H200 GPU (MarkTechPost). The release targets practitioners seeking efficiency gains through conditional computation on inference-constrained hardware.
Google has restructured its Gemini access tiers, restricting free users to Flash-Lite, the smallest available model, while Flash and Pro models are reserved for paid subscribers (The Decoder). The change, effective in October 2026, may signal preparation for launching Gemini 4 Argon, a more resource-intensive model (The Decoder).