Inference Cost
Curated collection of thoughts and builds centered around Inference Cost.
Daily Briefings 11
- Claude Code Auto Mode becomes default; agent energy consumption tracked at 600x chat baseline
- Alibaba Qwen3.8-Max reaches general availability; Thinking Machines releases 12B active MoE model
- DeepSeek V4 Flash Update Matches GPT-5.6 Luna at 60% Lower Cost; Thinking Machines Releases Smaller Inkling Model
- OpenAI cuts GPT-5.6 Luna pricing 80%; Google DeepMind releases Gemini Robotics 2 models
- Token Saver MCP cuts PDF costs 90-99%; Moonshot open-sources MoonEP for MoE training
- Cursor's Agent Swarm Achieves 100% SQLite Rebuild; Black Forest Labs Releases FLUX 3 Multimodal Model
- Anthropic releases Claude Opus 5 at unchanged pricing; Datalab Marker v2 benchmarks document processing
- Gigatoken tokenizer hits 24.5 GB/s; Cursor Router cuts inference costs 30–50%
- OpenAI models breach Hugging Face during internal security test; Google releases three new Gemini Flash variants
- Anthropic cuts Claude Fable limits; GPT-5.6 deletes files in full access mode
- GPT-5.6 Sol reported near Fable 5 at a third the cost; Kyutai ships open-weight MuScriptor