Inference
Curated collection of thoughts and builds centered around Inference.
Daily Briefings 19
- NVIDIA Releases NemotronLabs VoiceChat 11B Open Model; ByteDance Introduces SeedRealtime Multimodal LLM
- AMD Acquires Taalas for On-Chip Model Weights; Agent Plugin Standard Reaches 1.0
- AMD Open-Sources 16B MoE Model; NVIDIA Releases Molt Agentic RL Framework
- Liquid AI releases 8K-context encoders optimized for CPU inference; Fireworks launches routing layer for open-weight coding models
- Open Dreamer Ships Dreamer 4 Reproduction; Sakana AI Releases Fugu-Cyber Orchestration Model
- Poolside releases Laguna S 2.1 coding model; Runway launches model router for generative media
- Gigatoken tokenizer hits 24.5 GB/s; Cursor Router cuts inference costs 30–50%
- NVIDIA Releases Cosmos 3 Edge; Google's Frozen v2 Chip Targets 6-10x TPU Efficiency Gains
- Feyn Labs ships SQRL text-to-SQL family; Moonshot maxes GPU capacity on Kimi K3 in 48 hours
- Moonshot ships 2.8T Kimi K3 open weights; Hugging Face discloses AI-agent intrusion
- Google updates Gemma 4 with tool-calling fixes; xAI open-sources Grok-Build after data breach
- xAI ships Grok 4.5 at $2/M for coding agents; DeepSeek builds inference chip
- Fable 5 moves to usage-credit billing; Chinese open weights gain US enterprise traction
- OpenAI ships gpt-realtime-2.1 with 25% lower voice latency; Fable 5 moves to usage credits
- OpenAI previews GPT-5.6 Sol for approved partners; Anthropic Mythos export block eased
- OpenAI and Broadcom debut Jalapeño chip; GPT-5.6 previews three-tier model family
- DFlash block-diffusion decoding reports 15x Blackwell throughput; Mistral ships OCR 4
- OpenAI and Broadcom unveil Jalapeño inference chip targeting late-2026 deployment
- MiniMax M3 sparse-attention claims verified; Grok 4.3 lands on Amazon Bedrock