arXiv:2601.10315cs.CL2026-01

构建500条模拟法庭律师语音,用于测试语音识别系统辨别合成声音的能力。

ADVOSYNTH: A Synthetic Multi-Advocate Dataset for Speaker Identification in Courtroom Scenarios

  • 用Speech Llama Omni生成10个虚拟律师的合成语音。
  • 设计5组律师对辩场景,每组含2人共100条语音。
  • 适合研究语音合成检测与法庭语音识别的学者使用。

随着大规模语音到语音模型实现高保真度,结构化环境中合成语音的差异性成为重要研究课题。本文提出Advosynth-500数据集,包含100条合成语音文件,涵盖10个独特律师身份。基于Speech Llama Omni模型,模拟5组不同律师在法庭辩论中的互动。为每位律师设定特定声学特征,并构建语音识别挑战任务,评估现代系统将音频文件映射至其对应合成来源的能力。数据集已开源,地址:https://github.com/naturenurtureelite/ADVOSYNTH-500。

原文摘要 · Abstract (English)

As large-scale speech-to-speech models achieve high fidelity, the distinction between synthetic voices in structured environments becomes a vital area of study. This paper introduces Advosynth-500, a specialized dataset comprising 100 synthetic speech files featuring 10 unique advocate identities. Using the Speech Llama Omni model, we simulate five distinct advocate pairs engaged in courtroom arguments. We define specific vocal characteristics for each advocate and present a speaker identification challenge to evaluate the ability of modern systems to map audio files to their respective synthetic origins. Dataset is available at this link-https: //github.com/naturenurtureelite/ADVOSYNTH-500.

语音合成法庭语音说话人识别数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。