gekro
GitHub LinkedIn
News

AI News

Alibaba Qwen3.8-Max reaches general availability; Thinking Machines releases 12B active MoE model

Alibaba moved its 2.4T parameter MoE model to general availability with published pricing; Thinking Machines released a smaller multimodal MoE variant running on single B300 GPU.

1 min read 3 sources

Alibaba’s Qwen team transitioned Qwen3.8-Max from preview to general availability, with per-token pricing published and open weights scheduled for release next week (MarkTechPost). The 2.4 trillion parameter mixture-of-experts model accepts text, image, and video input across a 1 million token context window. The move adds a publicly priced option to the dense frontier model landscape, though no benchmark comparison table was published.

Thinking Machines Lab released Inkling-Small, a 276 billion total parameter multimodal MoE model with 12 billion active parameters that matches the full Inkling variant at a quarter the total size (MarkTechPost). The NVFP4 checkpoint runs on a single NVIDIA B300 GPU, reducing deployment and inference cost barriers for engineers targeting multimodal reasoning on constrained hardware. Cogent AI released VR-1, a reasoning model post-trained specifically for cybersecurity workflows, with IntrusionBench, a benchmark scoring agents on completed enterprise intrusions (MarkTechPost).

Compiled automatically from the linked sources and published without manual editing - a neutral summary of third-party reporting, for information only. Every claim links to its origin. Not original reporting.