arXiv:2505.02707cs.AIcs.CL2025-05被引 11

Voila让语音助手实时、主动、有情感地与人交互,支持自定义声音和多任务应用。

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play

  • 端到端架构实现低延迟全双工对话,响应仅195毫秒。
  • 支持百万级预设声音,10秒音频即可生成新声音。
  • 统一模型覆盖语音识别、合成与多语言翻译,适合语音交互研究者。

一个能无缝融入日常生活的语音智能体,应能以自主、实时且富有情感的方式与人类互动。它不应仅回应指令,而需持续聆听、推理并主动回应,促成流畅、动态且情感共鸣的交流。我们提出Voila,一系列大型语音-语言基础模型,朝着这一愿景迈出一步。Voila突破传统流水线系统,采用全新端到端架构,实现全双工、低延迟对话,同时保留丰富的语音细节,如语调、节奏与情感。其响应延迟仅为195毫秒,优于平均人类反应时间。其分层多尺度Transformer融合大语言模型的推理能力与强大的声学建模,实现自然、角色感知的语音生成——用户仅需撰写文本指令即可定义说话人身份、语气等特征。此外,Voila支持超过一百万种预设声音,并可从仅10秒的音频样本中高效定制新声音。除口语对话外,Voila被设计为统一模型,适用于自动语音识别(ASR)、文语转换(TTS),并通过少量调整支持多语言语音翻译。该模型完全开源,旨在推动开放研究,加速下一代人机交互的发展。

原文摘要 · Abstract (English)

A voice AI agent that blends seamlessly into daily life would interact with humans in an autonomous, real-time, and emotionally expressive manner. Rather than merely reacting to commands, it would continuously listen, reason, and respond proactively, fostering fluid, dynamic, and emotionally resonant interactions. We introduce Voila, a family of large voice-language foundation models that make a step towards this vision. Voila moves beyond traditional pipeline systems by adopting a new end-to-end architecture that enables full-duplex, low-latency conversations while preserving rich vocal nuances such as tone, rhythm, and emotion. It achieves a response latency of just 195 milliseconds, surpassing the average human response time. Its hierarchical multi-scale Transformer integrates the reasoning capabilities of large language models (LLMs) with powerful acoustic modeling, enabling natural, persona-aware voice generation -- where users can simply write text instructions to define the speaker's identity, tone, and other characteristics. Moreover, Voila supports over one million pre-built voices and efficient customization of new ones from brief audio samples as short as 10 seconds. Beyond spoken dialogue, Voila is designed as a unified model for a wide range of voice-based applications, including automatic speech recognition (ASR), Text-to-Speech (TTS), and, with minimal adaptation, multilingual speech translation. Voila is fully open-sourced to support open research and accelerate progress toward next-generation human-machine interactions.

语音交互大模型实时对话声音生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。