arXiv:2506.18296cs.SDeess.AS2025-06中稿 · on Interspeech 202…

构建日本偶像语音数据集,助力语音合成与风格迁移研究

JIS: A Speech Corpus of Japanese Idol Speakers with Various Speaking Styles

  • 聚焦日本年轻女性现场偶像,统一身份标签便于实验设计
  • 收录多位偶像真实发音,支持高保真语音生成与风格转换
  • 免费开放用于非商业研究,适合语音风格个性化方向探索

为推动语音生成人工智能研究(包括文本到语音合成TTS和语音转换VC),我们构建了日本偶像语音语料库(JIS)。JIS包含多位来自日本“年轻女性现场偶像”这一特定群体的说话者,每位均以舞台名标识,便于招募熟悉其声音的听众开展听觉实验。该语料库可实现对TTS和VC系统中说话人相似度的更严格评估。凭借其独特的说话人属性,JIS将推动声音定制化生成等尚未广泛研究的方向。语料库将免费提供,仅限非商业性基础研究使用。本文介绍了JIS的构建过程,概述了日本现场偶像文化背景以支持其伦理合规使用,并提供了基础分析以指导实际应用。

原文摘要 · Abstract (English)

We construct Japanese Idol Speech Corpus (JIS) to advance research in speech generation AI, including text-to-speech synthesis (TTS) and voice conversion (VC). JIS will facilitate more rigorous evaluations of speaker similarity in TTS and VC systems since all speakers in JIS belong to a highly specific category: "young female live idols" in Japan, and each speaker is identified by a stage name, enabling researchers to recruit listeners familiar with these idols for listening experiments. With its unique speaker attributes, JIS will foster compelling research, including generating voices tailored to listener preferences-an area not yet widely studied. JIS will be distributed free of charge to promote research in speech generation AI, with usage restricted to non-commercial, basic research. We describe the construction of JIS, provide an overview of Japanese live idol culture to support effective and ethical use of JIS, and offer a basic analysis to guide application of JIS.

语音合成数据集偶像语音TTS

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。