StepFun has released Step 5 Preview, a sparse Mixture-of-Experts model with 600B total parameters and 27B active parameters per token, supporting a 1M-token context window with multimodal input (text, image, video) (MarkTechPost). The model targets long-horizon agentic work in software engineering and professional knowledge tasks, addressing the emerging focus on token efficiency for agent orchestration seen across recent infrastructure updates.
Alibaba’s Qwen team released Qwen-Image-2.1, an open-weight image generation model with 7 billion parameters that claims to match closed-model performance and runs on consumer GPUs, supporting transparency and up to ten reference images (The Decoder). Additionally, Hugging Face published research on LLM pruning using physics-inspired Ising optimization for block removal (Hugging Face Blog), and Tencent introduced Gander, a multimodal agent architecture that maintains dialogue while executing background tasks via swappable “brain” modules (The Decoder).