Google Research has introduced WikiSkill, a framework that equips AI agents with a persistent knowledge base documenting both failures and successes in a wiki-like structure (The Decoder). Rather than discarding learned information after each run, agents retain and reference this knowledge to improve future performance. The approach shows larger models benefit most from accumulated experience, addressing a longstanding limitation in agentic systems where each task execution operates without institutional memory.
Anthropicannounced the Model Hardware Standard (MHS), a unified interface enabling AI agents to control physical devices such as robotic arms and lab instruments (The Decoder). Early tests show integration time dropped from weeks to hours, though Claude instances occasionally failed to grasp physical cause-and-effect relationships, indicating human oversight remains essential. Separately, LAION released its Big Video Dataset (BVD), containing 80 million videos totaling 10 million hours of runtime and 55 million auto-described clips; models trained on BVD outperformed the previous benchmark InternVid by up to 2.1 percentage points (The Decoder).