Open-weight decision models saw multiple releases in the past 36 hours, each targeting fast structured outputs without text generation. AWS Strands Labs released Strands Decider 2B, built on Qwen3.5-2B-Base and returning choices, probabilities, and confidence scores in a single forward pass, achieving a median latency of 115 ms on an RTX 3090 (MarkTechPost). Cloudflare released Clef (27B) and Clef-flash (9B), open-weight models that return typed probabilities instead of text, running on Workers AI at 209.3 ms and 38.8 ms median latency respectively (MarkTechPost). Both are Jev-API compatible and accept images.
Cohere released Embed 5, a new embedding model family designed for enterprise search, RAG, and agentic retrieval, shipped in two tiers: Embed 5 Pro for maximum retrieval quality and Embed 5 Fast for latency and cost optimization on live query paths, with support for text, images, and fused text-image inputs (MarkTechPost). Additionally, Hugging Face published a technical introduction to Olmo-core 3, described as open, scalable training infrastructure for large mixture-of-experts models (Hugging Face Blog).