arXiv:2509.08438cs.CLcs.MM2025-09

新数据集+生成框架,提升真实语音关系抽取效果

CommonVoice-SpeechRE and RPG-MoGe: Advancing Speech Relation Extraction with a New Dataset and Multi-Order Generative Framework

  • 采用多阶生成策略与关系提示引导跨模态对齐
  • 在近2万条真实语音上实现超越现有方法的性能
  • 适合语音信息抽取与多模态学习研究者使用

语音关系抽取(SpeechRE)旨在直接从语音中提取关系三元组。然而,现有基准数据集严重依赖合成数据,缺乏真实人类语音的数量与多样性。同时,现有模型受限于固定的单阶生成模板和弱语义对齐,显著制约性能。为此,我们提出CommonVoice-SpeechRE,一个包含近20,000条来自多样化说话人的真实语音样本的大规模数据集,为SpeechRE研究建立新基准。此外,我们提出关系提示引导的多阶生成集成框架RPG-MoGe:(1) 采用多阶三元组生成集成策略,在训练与推理阶段通过多样元素顺序利用数据多样性;(2) 使用基于CNN的潜在关系预测头,生成显式关系提示以引导跨模态对齐并实现精准三元组生成。实验表明,该方法优于当前最先进方法,既提供基准数据集,也给出适用于真实场景的高效解决方案。源代码与数据集已公开于https://github.com/NingJinzhong/SpeechRE_RPG_MoGe。

原文摘要 · Abstract (English)

Speech Relation Extraction (SpeechRE) aims to extract relation triplets directly from speech. However, existing benchmark datasets rely heavily on synthetic data, lacking sufficient quantity and diversity of real human speech. Moreover, existing models also suffer from rigid single-order generation templates and weak semantic alignment, substantially limiting their performance. To address these challenges, we introduce CommonVoice-SpeechRE, a large-scale dataset comprising nearly 20,000 real-human speech samples from diverse speakers, establishing a new benchmark for SpeechRE research. Furthermore, we propose the Relation Prompt-Guided Multi-Order Generative Ensemble (RPG-MoGe), a novel framework that features: (1) a multi-order triplet generation ensemble strategy, leveraging data diversity through diverse element orders during both training and inference, and (2) CNN-based latent relation prediction heads that generate explicit relation prompts to guide cross-modal alignment and accurate triplet generation. Experiments show our approach outperforms state-of-the-art methods, providing both a benchmark dataset and an effective solution for real-world SpeechRE. The source code and dataset are publicly available at https://github.com/NingJinzhong/SpeechRE_RPG_MoGe.

语音抽取多模态生成框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。