gekro
GitHub LinkedIn
News

AI News

AMD Acquires Taalas for On-Chip Model Weights; Agent Plugin Standard Reaches 1.0

AMD acquired Taalas to embed model weights directly in inference silicon, achieving 16K tokens/sec per user; five major companies jointly released Agent Plugins 1.0 standard.

1 min read 4 sources

AMD has acquired Taalas, a Canadian startup that hard-codes AI model weights directly into inference chips, according to reporting from (The Decoder). The approach trades flexibility for speed; a demo chip running Llama 3.1-8B achieved over 16,000 tokens per second per user, though each chip is locked to a single model. (The Decoder) reports that Google is pursuing a similar strategy.

In separate infrastructure developments, Amazon, Cursor, Microsoft, OpenAI, and Vercel have jointly created Agent Plugins, an open standard defining a single package format for AI agent extensions, with version 1.0.0 released using a plugin.json manifest file and supporting both agent skills and MCP servers (The Decoder). Tencent Cloud open-sourced TencentDB Agent Memory v2.0, an MIT-licensed self-hosted memory hub for AI coding agents that transforms conversations, documents, and code into reusable Chat Memory, Skill, LLM-Wiki, and Code-Graph assets (MarkTechPost). NVIDIA Labs released NOOA, a model-agnostic Python framework that consolidates agent development into a single Python class (MarkTechPost).

Compiled automatically from the linked sources and published without manual editing - a neutral summary of third-party reporting, for information only. Every claim links to its origin. Not original reporting.