arXiv:2602.07036cs.SDcs.AI2026-02被引 3

构建多语种角色对话数据集,助力语音大模型训练

MENASpeechBank: A Reference Voice Bank with Persona-Conditioned Multi-Turn Conversations for AudioLLMs

  • 用角色设定+场景匹配生成41.7万条合成对话
  • 覆盖124位来自MENA地区的真实发音人,共1.8万条语音
  • 适合语音大模型、多语种对话系统研究者使用

语音大语言模型(AudioLLMs)能够对语音和通用音频进行指令理解,但进展日益受限于缺乏多样、对话式且与指令对齐的语音-文本数据。这一瓶颈在角色化交互和方言覆盖方面尤为突出,因真实多说话人录音的采集与发布成本高、耗时长。本文提出MENASpeechBank,一个包含约1.8万条高质量语音、来自124位跨中东与北非(MENA)国家说话人的参考语音库,涵盖英语、现代标准阿拉伯语(MSA)及地方阿拉伯语变体。基于该资源,我们开发了一套可控合成数据流程:(i)构建融合世界价值观调查特征的角色档案;(ii)定义约5000个对话场景分类;(iii)通过语义相似度匹配角色与场景;(iv)利用大语言模型生成约41.7万条角色扮演对话,用户以角色身份发言,助手作为助人代理响应;(v)通过参考语音音频条件化合成用户发言,保留说话人身份与多样性。我们评估了合成与真人录音的对话,并提供详细分析。我们将公开MENASpeechBank及生成对话数据集供社区使用。

原文摘要 · Abstract (English)

Audio large language models (AudioLLMs) enable instruction-following over speech and general audio, but progress is increasingly limited by the lack of diverse, conversational, instruction-aligned speech-text data. This bottleneck is especially acute for persona-grounded interactions and dialectal coverage, where collecting and releasing real multi-speaker recordings is costly and slow. We introduce MENASpeechBank, a reference speech bank comprising about 18K high-quality utterances from 124 speakers spanning multiple MENA countries, covering English, Modern Standard Arabic (MSA), and regional Arabic varieties. Building on this resource, we develop a controllable synthetic data pipeline that: (i) constructs persona profiles enriched with World Values Survey-inspired attributes, (ii) defines a taxonomy of about 5K conversational scenarios, (iii) matches personas to scenarios via semantic similarity, (iv) generates about 417K role-play conversations with an LLM where the user speaks as the persona and the assistant behaves as a helpful agent, and (v) synthesizes the user turns by conditioning on reference speaker audio to preserve speaker identity and diversity. We evaluate both synthetic and human-recorded conversations and provide detailed analysis. We will release MENASpeechBank and the generated conversations publicly for the community.

语音大模型角色对话多语种合成数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。