gekro
GitHub LinkedIn
News

AI News

NVIDIA Releases NemotronLabs VoiceChat 11B Open Model; ByteDance Introduces SeedRealtime Multimodal LLM

NVIDIA and ByteDance each released open-weight multimodal models for real-time audio-visual interaction with sub-500ms latency and native tool calling.

1 min read 3 sources

NVIDIA released NemotronLabs VoiceChat 11B, an open full-duplex speech-to-speech model achieving approximately 450 milliseconds turn-taking latency with live tool calling capability (MarkTechPost). The 11B parameter model enables continuous audio interaction without per-turn latency penalties typical of serialized pipelines, addressing a core bottleneck in conversational AI deployment.

ByteDance’s Seed team introduced SeedRealtime, a native audio-visual full-duplex LLM that fuses audio, video, and text in a unified architecture (MarkTechPost). Unlike models processing discrete turns, SeedRealtime operates over continuous multimodal streams, positioning it as a step toward omnichannel interaction patterns. Google DeepMind also released DiffusionGemma, which retrofitted Gemma 4 into a diffusion model using less than 10 percent of original training compute, achieving approximately 1,500 tokens per second through parallel decoding instead of sequential generation (The Decoder).

Compiled automatically from the linked sources and published without manual editing - a neutral summary of third-party reporting, for information only. Every claim links to its origin. Not original reporting.