Benchmarks
Curated collection of thoughts and builds centered around Benchmarks.
Daily Briefings 9
- Feyn Labs ships SQRL text-to-SQL family; Moonshot maxes GPU capacity on Kimi K3 in 48 hours
- Fable 5 moves to usage-credit billing; Chinese open weights gain US enterprise traction
- OpenAI ships gpt-realtime-2.1 with 25% lower voice latency; Fable 5 moves to usage credits
- Z.ai's GLM-5.2 beats Claude Code on Semgrep security benchmarks; MCP beta SDKs drop
- Z.ai's GLM-5.2 leads open-weight models with 1M context under MIT; Majors says code economics flipped
- NVIDIA releases 550B Nemotron 3 Ultra; Anthropic files for IPO and splits agent billing
- MiniMax M3 open-weight model claims SWE-Bench Pro lead with 1M-token context
- Microsoft unveils first in-house reasoning model; Anthropic scales Glasswing to 150 orgs
- GitHub Copilot shifts to token-based billing; transformer weather model matches ECMWF