arXiv:2505.13338cs.CLcs.AI2025-05中稿 · Interspeech 2025被引 5

构建首个融合语境与语气理解的语音问答数据集,提升语音大模型推理能力。

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation

  • 基于伪语气标签压缩真实语音数据,生成紧凑语料。
  • 利用大模型自动生成语境+语气联合问答对,效果接近人工标注。
  • 揭示语音大模型在共情推理上的短板,适合语音多模态研究者使用。

当前语音大模型在语境推理与语气理解方面能力有限,主要因缺乏同时涵盖两者的问答数据集。本文提出一种从真实语音数据中生成数据的新框架,将语境推理与语气信息相结合。该框架包含基于伪语气标签的数据压缩方法和基于大模型的上下文语气问答(CPQA)生成。在Qwen2-Audio-7B-Instruct模型上评估显示,本框架生成的数据与人工标注数据表现高度相关。结果还揭示语音大模型在共情推理任务中的不足,凸显此类数据集的重要性及更鲁棒模型的必要性。该框架为首个同类方案,具备训练更强语气推理语音大模型的潜力。

原文摘要 · Abstract (English)

Current speech-LLMs exhibit limited capability in contextual reasoning alongside paralinguistic understanding, primarily due to the lack of Question-Answer (QA) datasets that cover both aspects. We propose a novel framework for dataset generation from in-the-wild speech data, that integrates contextual reasoning with paralinguistic information. It consists of a pseudo paralinguistic label-based data condensation of in-the-wild speech and LLM-based Contextual Paralinguistic QA (CPQA) generation. The effectiveness is validated by a strong correlation in evaluations of the Qwen2-Audio-7B-Instruct model on a dataset created by our framework and human-generated CPQA dataset. The results also reveal the speech-LLM's limitations in handling empathetic reasoning tasks, highlighting the need for such datasets and more robust models. The proposed framework is first of its kind and has potential in training more robust speech-LLMs with paralinguistic reasoning capabilities.

语音大模型多模态数据生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。