构建细粒度语音风格数据集,实现跨粒度语音文本统一表征。
Towards Fine-Grained and Multi-Granular Contrastive Language-Speech Pre-training
- 基于端到端管道构建47小时语音与1900万条细粒度描述数据集
- 提出CLSP模型,在多粒度检索与风格评分中表现优异
- 适合语音风格分析与跨模态理解研究者使用
细粒度说话风格建模在语言-语音表征预训练中仍具挑战性,因现有模型通常依赖粗粒度标题或任务特定监督,且缺乏可扩展的细粒度风格标注。本文提出FCaps,一个包含47,000小时语音与1900万条细粒度自由文本风格描述的大规模数据集,通过新型端到端管道直接将详细描述与音频对齐,避免了现有级联流程中基于大模型重写带来的误差传播。基于LLM作为裁判的评估表明,其标注在准确性、覆盖率和自然度上均优于现有级联标注。在此基础上,我们提出CLSP,一种融合全局与细粒度监督的对比语言-语音预训练模型,实现多粒度统一表征。大量实验表明,CLSP学习到的细粒度、多粒度语音-文本表征在全局与细粒度语音-文本检索、零样本副语言分类及语音风格相似性评分任务中表现稳健,且与人类判断高度一致。代码与数据集已公开于https://github.com/yfyeung/CLSP。
原文摘要 · Abstract (English)
Modeling fine-grained speaking styles remains challenging for language-speech representation pre-training, as existing speech-text models are typically trained with coarse captions or task-specific supervision, and scalable fine-grained style annotations are unavailable. We present FCaps, a large-scale dataset with fine-grained free-text style descriptions, encompassing 47k hours of speech and 19M fine-grained captions annotated via a novel end-to-end pipeline that directly grounds detailed captions in audio, thereby avoiding the error propagation caused by LLM-based rewriting in existing cascaded pipelines. Evaluations using LLM-as-a-judge demonstrate that our annotations surpass existing cascaded annotations in terms of correctness, coverage, and naturalness. Building on FCaps, we propose CLSP, a contrastive language-speech pre-trained model that integrates global and fine-grained supervision, enabling unified representations across multiple granularities. Extensive experiments demonstrate that CLSP learns fine-grained and multi-granular speech-text representations that perform reliably across global and fine-grained speech-text retrieval, zero-shot paralinguistic classification, and speech style similarity scoring, with strong alignment to human judgments. Code and dataset are publicly available at https://github.com/yfyeung/CLSP.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。