gekro
GitHub LinkedIn
AI News · Updated daily

AI News

RSS · 92 briefings

Daily AI industry briefings - model releases, research, infrastructure, and pricing - summarized neutrally from vetted sources, with every claim linked to its origin.

  1. Latest

    Nvidia acquires Hugging Face for $12.9B; GPT-6 Astra benchmarks diverge while Astra shifts to on-device compute routing

    Nvidia confirmed acquisition of Hugging Face for $12.9 billion; separately, GPT-6 Astra showed conflicting benchmark results and Nvidia launched PAIR to route local AI tasks across home networks.

    TechCrunch The Decoder MarkTechPost
    Read →
  2. Perplexity open-sources Lily inference engine; Qwen releases local search layer

    Perplexity released Lily, a Rust-Metal inference engine for Apple Silicon showing 1.35x decode speedup over MLX-LM; Qwen open-sourced zg, a unified local-first search combining ripgrep, BM25, and.

    MarkTechPost MIT Technology Review
    Read →
  3. Google's Gemini cuts video token usage 88%; Anthropic releases Fable 5.1 with 45% cost reduction

    Google deploys agent-based video analysis to reduce Gemini token consumption; Anthropic releases Claude Fable 5.1 with improved coding and lower inference costs.

    The Decoder MarkTechPost
    Read →
  4. OpenAI expands outcome-based pricing; China's CXMT reaches HBM3E production

    OpenAI begins charging select customers only when tasks complete successfully, while China's memory maker produces high-bandwidth memory for AI chips.

    The Decoder
    Read →
  5. OpenAI introduces outcome-based pricing; EU classifies ChatGPT as very large search engine

    OpenAI is piloting outcome-based pricing with large customers, while the EU Commission designates ChatGPT as a very large online platform subject to stricter DSA oversight by year-end.

    The Decoder The Verge
    Read →
  6. Google introduces WikiSkill agent memory framework; Anthropic unveils Model Hardware Standard for robotics

    Google's WikiSkill gives AI agents persistent memory to learn from past failures and successes. Anthropic releases Model Hardware Standard to simplify agent integration with physical devices.

    The Decoder
    Read →
  7. GLM-5.3-Flash and Qwen3.8-Flash converge on identical architecture; Google DeepMind Co-Scientist automates lab work

    Two independent Chinese labs shipped functionally identical model architectures; Google's AI system now plans experiments, operates equipment, and writes papers.

    MarkTechPost The Decoder
    Read →
  8. Google's Gemini Omni 1.1 Flash extends video generation; agent sandbox pricing comparison emerges

    Google reduces video generation costs and latency with Gemini Omni 1.1 Flash; developer tools for agent execution environments face fragmented pricing models.

    The Decoder Google DeepMind's technical blog MarkTechPost
    Read →
  9. OpenAI's Jalapeño inference chip outperforms Nvidia on throughput and efficiency; IBM releases Granite 4.2 open-weight models

    OpenAI's custom Jalapeño chip demonstrated superior inference throughput and energy efficiency versus Nvidia's latest offerings, while IBM published technical details on Granite 4.2 LLMs.

    The Decoder TechCrunch Hugging Face Blog
    Read →
  10. Cerebras CS-4 doubles performance on same chip; Alibaba's Wan3.0 video model reaches 30-second generation

    Cerebras released the CS-4 accelerator with double the performance of its predecessor; Alibaba launched Wan3.0 for text-to-video generation up to 30 seconds.

    The Decoder Hugging Face Blog
    Read →
  11. FreeToken runs 753B MoE models on single GPU; agentic token consumption surges 14x

    Local inference engine enables frontier models on consumer hardware; agent-driven workloads now dominate token spend on OpenRouter.

    MarkTechPost The Decoder
    Read →
  12. Agent loop harness engineering outweighs model choice; safety benchmarks show structural flaws

    Research shows agent performance depends more on orchestration architecture than base model; psychological analysis reveals safety benchmarks don't measure consistent traits.

    MarkTechPost TechCrunch The Decoder
    Read →
  13. DeepSeek V4-Flash-Vision rivals Opus 4.8 on agent benchmarks; safety testing reveals benchmark gaming

    DeepSeek releases experimental multimodal model approaching frontier performance; UK researchers find popular safety benchmarks don't measure consistent traits and can be artificially inflated.

    The Decoder
    Read →
  14. OpenAI patches Codex file-deletion bug; Anthropic demonstrates agent-driven protein design

    OpenAI fixed a critical Codex bug that deleted user files without permission; Anthropic showed Claude agents designing proteins at 35% hit rate via tool orchestration.

    The Decoder
    Read →
  15. NVIDIA TensorRT Model Connect enables two-command Hugging Face to C++ inference; Google open-sources SAM agent mesh for MCP discovery

    NVIDIA released TensorRT Model Connect for direct Hugging Face checkpoint compilation; Google open-sourced SAM, a zero-config P2P network for agent MCP tool discovery across environments.

    MarkTechPost The Decoder
    Read →
  16. AWS CPU strain prompts infrastructure rethink; OpenAI dissolves preparedness team

    AWS faces CPU capacity bottlenecks from AI workloads, spurring conservation mandates; OpenAI shut down its team evaluating catastrophic model risks.

    IEEE Spectrum The Decoder
    Read →
  17. Optima benchmarking platform lets engineers test models on custom data; Anthropic reveals year-long safety filter outage

    Artificial Analysis launched Optima for custom AI benchmarking; Anthropic disclosed its bio-weapons filter was inactive for nearly a year, exposing 133 million unfiltered requests.

    The Decoder
    Read →
  18. Alibaba Qwen 3.8 open-weights release; GLM-5.3 claims strongest coding model via post-training

    Alibaba released Qwen 3.8 (27B, Apache 2.0) with 262K context; Zhipu AI's GLM-5.3 shows 50% coding gains without base retraining.

    The Decoder MarkTechPost
    Read →
  19. Google Gemini 3.7 Flash undercuts predecessor 50%; DeepSeek open-sources agent harness

    Google released Gemini 3.7 Flash with improved coding performance at half the price of its three-week-old predecessor; DeepSeek shipped V4 Pro and open-sourced Harness agent software under MIT.

    The Decoder TechCrunch
    Read →
  20. SpaceXAI's Grok 4.6 matches frontier performance at lower cost; Dyna-2 scales robot learning to 1M video hours

    Grok 4.6 ties top models on benchmarks while undercutting price; Dyna Robotics releases world-action model trained on million hours of egocentric video.

    MarkTechPost The Decoder
    Read →
  21. NVIDIA Releases Nemotron 3.5 Lightning MoE and LTX-2.5 Open Video Model

    NVIDIA released two open-weight models targeting inference efficiency: Nemotron 3.5 Lightning, a 30B MoE with 3B active parameters for agent execution, and LTX-2.5, a world model for local video.

    MarkTechPost The Decoder
    Read →
  22. Meta releases Muse Glimmer 30B agentic model; webAI open-sources formal-logic models for local inference

    Meta released Muse Glimmer, a 30B open-weight agentic model running on consumer GPUs; webAI shipped 1.7B and 3B formal-logic models for on-device reasoning.

    MarkTechPost The Decoder
    Read →
  23. NVIDIA Releases NemotronLabs VoiceChat 11B Open Model; ByteDance Introduces SeedRealtime Multimodal LLM

    NVIDIA and ByteDance each released open-weight multimodal models for real-time audio-visual interaction with sub-500ms latency and native tool calling.

    MarkTechPost The Decoder
    Read →
  24. Claude Code Auto Mode becomes default; agent energy consumption tracked at 600x chat baseline

    Anthropic defaults Claude Code to Auto Mode for safer command approval starting August 14; energy researcher measures agentic workloads at 600 times the energy cost of standard chat.

    The Decoder TechCrunch
    Read →
  25. AMD Acquires Taalas for On-Chip Model Weights; Agent Plugin Standard Reaches 1.0

    AMD acquired Taalas to embed model weights directly in inference silicon, achieving 16K tokens/sec per user; five major companies jointly released Agent Plugins 1.0 standard.

    The Decoder MarkTechPost
    Read →
  26. Liquid AI Releases On-Device Agentic Model; Microsoft Open-Sources Unit-Test Agent

    Liquid AI released LFM2.5-2.6B, a 2.69B parameter agentic model with 128K context and on-device tool calling; Microsoft open-sourced code-testing-generator, a polyglot unit-test agent achieving.

    MarkTechPost
    Read →
  27. Meta launches Muse Code agent for large codebases; Mistral's 3B Shieldstral matches larger safety models

    Meta released Muse Code, an AI agent for complex software tasks, while Mistral's compact 3B safety model matches systems seven times its size.

    TechCrunch The Decoder
    Read →
  28. NVIDIA Releases Alpamayo 2 Super Open Vision-Language-Action Model; CopilotKit Open-Sources Channels SDK for Agent Deployment

    NVIDIA released a 34B open-weight vision-language-action model for autonomous driving; CopilotKit published an MIT-licensed SDK for deploying agents in Slack and Teams.

    MarkTechPost
    Read →
  29. Y Combinator open-sources QM multiplayer agent harness; MiniMax H3 becomes first open model to top video ranking

    Y Combinator released QM, a multiplayer agent framework for Slack and web; MiniMax open-sourced H3 video model, ranking atop video benchmarks.

    MarkTechPost The Decoder
    Read →
  30. Alibaba Qwen3.8-Max reaches general availability; Thinking Machines releases 12B active MoE model

    Alibaba moved its 2.4T parameter MoE model to general availability with published pricing; Thinking Machines released a smaller multimodal MoE variant running on single B300 GPU.

    MarkTechPost
    Read →

Summarized from vetted sources, every claim linked. For information only.