跨语言语音与文本联合建模,提升隐含语篇关系识别效果
Multilingual Extraction and Recognition of Implicit Discourse Relations in Speech and Text
- 构建英法西三语多模态数据集,融合文本与语音特征
- 音视频联合建模使低资源语言性能提升显著
- 适用于多语言语篇分析、跨语言对话系统研发
隐含语篇关系分类是一项挑战性任务,需从上下文中推断语义。由于语境线索可能分布在不同模态且跨语言差异大,仅依赖文本常无法充分捕捉信息。为此,我们提出一种自动方法,针对远相关和无关语言对,构建涵盖英语、法语和西班牙语的多语言多模态数据集。在分类方面,提出基于Qwen2-Audio的多模态方法,实现文本与音频的联合建模,以支持跨语言的隐含语篇关系识别。实验发现,尽管纯文本模型表现优于纯音频模型,但融合双模态可提升性能,且跨语言迁移对低资源语言有显著增益。
原文摘要 · Abstract (English)
Implicit discourse relation classification is a challenging task, as it requires inferring meaning from context. While contextual cues can be distributed across modalities and vary across languages, they are not always captured by text alone. To address this, we introduce an automatic method for distantly related and unrelated language pairs to construct a multilingual and multimodal dataset for implicit discourse relations in English, French, and Spanish. For classification, we propose a multimodal approach that integrates textual and acoustic information through Qwen2-Audio, allowing joint modeling of text and audio for implicit discourse relation classification across languages. We find that while text-based models outperform audio-based models, integrating both modalities can enhance performance, and cross-lingual transfer can provide substantial improvements for low-resource languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。