gekro
GitHub LinkedIn
News

AI News

FreeToken runs 753B MoE models on single GPU; agentic token consumption surges 14x

Local inference engine enables frontier models on consumer hardware; agent-driven workloads now dominate token spend on OpenRouter.

1 min read 4 sources

Researchers introduced FreeToken, an edge-native serving engine that runs a 753 billion parameter mixture-of-experts model on a single workstation GPU by splitting MoE cache misses between PCIe and CPU execution based on measured bandwidths (MarkTechPost). The approach demonstrates a path toward running frontier-scale models locally without cloud infrastructure, addressing the hardware constraints that typically limit on-device inference.

Token economics shifted significantly as agentic workloads surpassed human usage on OpenRouter starting February 2025, with agent consumption growing 14x since then compared to 2.8x growth in human usage (The Decoder). Approximately 70 percent of agent token spend comes from cached prompts, meaning actual cost growth trails token growth substantially. Meanwhile, Harvey introduced Harvey Tenet, a post-trained model based on Kimi K3 tuned via Fireworks for legal agent tasks, though independent verification of benchmark claims remains limited (MarkTechPost). Vercel and Ora released Is Agentic, a free audit tool that scores website readiness for agent interaction across 118 checks (MarkTechPost).

Compiled automatically from the linked sources and published without manual editing - a neutral summary of third-party reporting, for information only. Every claim links to its origin. Not original reporting.