gekro
GitHub LinkedIn
News

AI News

StepFun releases Step 5 Preview 600B MoE model; Alibaba open-sources Qwen-Image-2.1 with 7B parameters

StepFun launched Step 5 Preview, a sparse MoE model with 1M-token context for agentic work; Alibaba released open-weight Qwen-Image-2.1 for image generation on consumer GPUs.

1 min read 4 sources

StepFun has released Step 5 Preview, a sparse Mixture-of-Experts model with 600B total parameters and 27B active parameters per token, supporting a 1M-token context window with multimodal input (text, image, video) (MarkTechPost). The model targets long-horizon agentic work in software engineering and professional knowledge tasks, addressing the emerging focus on token efficiency for agent orchestration seen across recent infrastructure updates.

Alibaba’s Qwen team released Qwen-Image-2.1, an open-weight image generation model with 7 billion parameters that claims to match closed-model performance and runs on consumer GPUs, supporting transparency and up to ten reference images (The Decoder). Additionally, Hugging Face published research on LLM pruning using physics-inspired Ising optimization for block removal (Hugging Face Blog), and Tencent introduced Gander, a multimodal agent architecture that maintains dialogue while executing background tasks via swappable “brain” modules (The Decoder).

Compiled automatically from the linked sources and published without manual editing - a neutral summary of third-party reporting, for information only. Every claim links to its origin. Not original reporting.