NVIDIA released NemotronLabs VoiceChat 11B, an open full-duplex speech-to-speech model achieving approximately 450 milliseconds turn-taking latency with live tool calling capability (MarkTechPost). The 11B parameter model enables continuous audio interaction without per-turn latency penalties typical of serialized pipelines, addressing a core bottleneck in conversational AI deployment.
ByteDance’s Seed team introduced SeedRealtime, a native audio-visual full-duplex LLM that fuses audio, video, and text in a unified architecture (MarkTechPost). Unlike models processing discrete turns, SeedRealtime operates over continuous multimodal streams, positioning it as a step toward omnichannel interaction patterns. Google DeepMind also released DiffusionGemma, which retrofitted Gemma 4 into a diffusion model using less than 10 percent of original training compute, achieving approximately 1,500 tokens per second through parallel decoding instead of sequential generation (The Decoder).