How to Build and Scale Voice Agents Using NVIDIA Nemotron, Modal, and Daily | Nemotron Labs

How to Build and Scale Voice Agents Using NVIDIA Nemotron, Modal, and Daily | Nemotron Labs

🎙 NVIDIA Developer 👥 222K 📅 March 4, 2026 ⏱ 53 min 👁 5K 📄 tutorial 🧭 2026-08-13
Available in: English (current) Français

Keywords

voice agentsNVIDIA NemotronModalDailyreal-time AI

Summary

This livestream from NVIDIA Developer demonstrates how to build and scale real-time voice agents using NVIDIA’s open Nemotron models, orchestrated on Modal and powered by Daily for real-time audio and agent communication. The session begins with a live demo of a voice agent that uses speech-to-text, LLM, and text-to-speech services running on NIM containers, deployed on Modal, and connected via Daily’s WebRTC infrastructure. The hosts then walk through the code architecture, showing how to set up each service (ASR, LLM, TTS) as separate Modal apps with GPU acceleration, and how to use PipeCat to orchestrate the pipeline. They discuss the importance of low latency and concurrency, and compare cascaded pipelines (STT -> LLM -> TTS) with emerging speech-to-speech models, highlighting the trade-offs in terms of latency, intelligence, and cost. The session also includes a second demo from CES featuring a robot controlled by a voice agent using a DGX Spark. The hosts answer audience questions about scaling, model selection, and the future of voice AI, emphasizing the benefits of using multiple models for different tasks.

175 words

Critical Evaluation

Value of the Information & Strength of the Argument

The video provides high practical value for developers interested in building voice agents, offering a step-by-step guide with real code and live demonstrations. The argumentation is solid, grounded in hands-on experience and technical details. The hosts explain the reasoning behind architectural choices, such as running each model as a separate service to allow independent scaling and avoid GPU contention. They also provide a balanced comparison between cascaded pipelines and speech-to-speech models, acknowledging the potential of the latter while justifying the current preference for the former in production due to latency, instruction-following accuracy, and cost. The discussion is well-structured and addresses common concerns, such as information loss in STT and the trade-offs of different model architectures.

Scientific Rigor, Source Quality, Title Accuracy

The scientific rigor is high for a technical tutorial: the hosts demonstrate real working systems and reference open-source projects and documentation. The sources cited are primarily official documentation and repositories from NVIDIA, Modal, and Daily, which are relevant and credible. The title accurately reflects the content, which is a practical guide to building and scaling voice agents. The video does not include formal citations, but the practical demonstrations and references to open-source code provide a solid foundation. The audience comments are not provided, so no analysis of public reception is included.

221 words

Title / Content Match

The title accurately reflects the content, which focuses on building and scaling voice agents using NVIDIA Nemotron, Modal, and Daily.

Quality & Reliability

8/10

High technical quality with live demos and code walkthroughs from industry experts. Claims are supported by practical examples and references to open-source projects. Minor limitations: no formal citations, some claims about model performance are anecdotal.

Key Moments

Cited Sources

Concurring Sources

Contribution & Novelties

The video provides a practical, hands-on guide to building and scaling voice agents using a combination of NVIDIA Nemotron models, Modal’s serverless infrastructure, and Daily’s real-time communication platform. It demonstrates a working system with live demos and code walkthroughs, offering insights into architectural patterns for low-latency and high-concurrency voice AI. The discussion on cascaded pipelines vs. speech-to-speech models provides a nuanced perspective on current production trade-offs. The video also highlights the importance of using multiple models for different tasks and the benefits of open-source models for self-hosting.

Pour aller plus loin :

  • WebRTC — Core technology for real-time audio/video communication used in the demo.
  • NVIDIA NIM — Containerized microservices for AI models, central to the deployment.
  • Serverless Computing — Concept underlying Modal’s infrastructure.
  • Speech-to-Text — Key component of the voice agent pipeline.
  • Text-to-Speech — Key component of the voice agent pipeline.

141 words

Radar Profile

The radar profile shows high scores across all dimensions, indicating a well-rounded technical tutorial with strong information content, quality, technical depth, and reliability. The video excels in providing practical, actionable knowledge.

Reliability 8/10