gekro
GitHub LinkedIn
News

AI News

Nvidia's SoL-Pi cuts coding agent token usage by 49%; Liquid AI releases speculative decoding for vision-language models

Nvidia's SoL-Pi system reduces coding agent token consumption by up to 49 percent through control layer optimization; Liquid AI ships speculative decoding draft model delivering up to 3.13x faster.

1 min read 3 sources

Nvidia’s SoL-Pi system optimizes the control harness between language models and their coding environments to reduce token usage by up to 49 percent with minimal performance degradation. The research involved testing 152 different approaches across more than 3,000 runs to develop the system, though gains were smaller on other benchmarks. (The Decoder)

Liquid AI released LFM2.5-VL-3B-DSpark, a 279.5M-parameter draft model enabling speculative decoding for its LFM2.5-VL-3B vision-language model. The draft model achieves up to 3.13x faster decoding on Apple M5 Max and 2.66x on H100 hardware with identical output under greedy decoding, with support shipping in llama.cpp and other inference frameworks. (MarkTechPost) Perplexity published a post-training study showing its Computer agent reduced tool-call failures from 2.24 percent to 1.77 percent by training on real user sessions paired with hint-guided self-distillation, demonstrating improvements through learning from mistakes in live deployment. (MarkTechPost)

Compiled automatically from the linked sources and published without manual editing - a neutral summary of third-party reporting, for information only. Every claim links to its origin. Not original reporting.