arXiv:2608.25177cs.SDcs.AI2026-08

用自然语言指令灵活聚类语音,支持多视角分析。

AudioLens: Multi-Perspective Speech Clustering with Reasoning Audio-Language Models

论文配图:AudioLens: Multi-Perspective Speech Clustering with Reasoning Audio-Language Models
图 1 · 摘自论文原文
  • 直接根据自然语言指令聚类语音,自动确定簇数和分配。
  • 在多个领域数据集上提升12.99点ARI和11.62点V-measure。
  • 适合需要按语义或语气重新组织语音库的研究者。

语音聚类是组织快速增长语音数据的基础任务,支持对话分析和语音驱动发现等应用。然而,现有方法依赖固定声学相似性度量或基于ASR的文本管道,难以根据不同用户指定的视角重新组织同一语音集合,尤其当聚类需同时考虑语言和副语言线索时。我们提出音频多视角聚类,即模型能根据自然语言视角直接划分语音记录,并推断簇数与归属。为此,我们构建了AudioLens-Bench基准,覆盖多个应用领域,评估视角内与跨视角泛化能力。进一步提出AudioLens-R1,一个端到端的大规模音频-语言模型,通过推理蒸馏与偏好优化训练。实验表明,AudioLens-R1持续优于所有基线,在整体ARI上提升12.99点,V-measure提升11.62点。结果证明原生音频-语言模型在语音集合的灵活、视角条件化结构发现中的潜力。

原文摘要 · Abstract (English)

Audio clustering is a fundamental task for organizing rapidly growing speech collections, supporting applications such as conversational analysis and speech-driven discovery. However, existing methods rely on fixed acoustic similarity metrics or ASR-based text pipelines, limiting their ability to reorganize the same audio collection under different user-specified perspectives, especially when clustering depends on both linguistic and paralinguistic cues. We introduce audio multi-perspective clustering, where a model directly partitions speech recordings according to a natural-language perspective while inferring both the number of clusters and their assignments. To study this setting, we construct AudioLens-Bench, a benchmark spanning multiple application domains and evaluating both in-perspective and cross-perspective generalization. We further propose AudioLens-R1, an end-to-end large audio-language model trained with reasoning distillation and preference optimization. Experiments show that AudioLens-R1 consistently outperforms all baselines, improving overall ARI by 12.99 points and V-measure by 11.62 points. These results demonstrate the promise of native audio-language models for flexible, perspective-conditioned structure discovery over speech collections.

语音聚类多模态大模型自然语言指令

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。