Alibaba’s Qwen team transitioned Qwen3.8-Max from preview to general availability, with per-token pricing published and open weights scheduled for release next week (MarkTechPost). The 2.4 trillion parameter mixture-of-experts model accepts text, image, and video input across a 1 million token context window. The move adds a publicly priced option to the dense frontier model landscape, though no benchmark comparison table was published.
Thinking Machines Lab released Inkling-Small, a 276 billion total parameter multimodal MoE model with 12 billion active parameters that matches the full Inkling variant at a quarter the total size (MarkTechPost). The NVFP4 checkpoint runs on a single NVIDIA B300 GPU, reducing deployment and inference cost barriers for engineers targeting multimodal reasoning on constrained hardware. Cogent AI released VR-1, a reasoning model post-trained specifically for cybersecurity workflows, with IntrusionBench, a benchmark scoring agents on completed enterprise intrusions (MarkTechPost).