Data from Andreessen Horowitz shows a Jevons paradox emerging in AI infrastructure: as token prices continue to fall, demand grows faster than costs drop, keeping H100 GPU rental prices steady or rising (The Decoder). The dynamic means that cheaper inference does not automatically reduce total spending by cloud providers and chip makers - instead, it unlocks new workloads faster than efficiency gains can offset hardware demand. This pattern holds until demand growth flattens, creating a potential inflection point for the infrastructure stack.
Microsoft released Decision-1, a 9B model built on Qwen3.5-9B and optimized for classification, routing, and agent control, achieving 83.5 percent accuracy with 85 ms latency across 36 benchmarks (The Decoder, MarkTechPost). The model returns calibrated probabilities for fixed answer options rather than generated text, positioning it for agent routing and orchestration workflows. Concurrently, Sakana AI published results showing a Claude-based multi-agent peer review system catching 73.43 percent of core-claim errors in scientific papers versus 14.81 percent for prior methods (MarkTechPost).