arXiv:2412.04026cs.CL2024-12被引 6

构建多模态多语言文档级信息抽取数据集M³D,支持中英文视频文本联合分析。

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction

  • 构建含文档级图文/视频对的多语言数据集,支持实体、关系、链式抽取与视觉定位
  • 在英中文数据上平均准确率达53.80%和53.77%,验证了模型有效性
  • 引入传记主题并设计抗缺失模态模块,适合跨语言多模态研究者使用

多模态信息抽取(IE)因融合多源信息提升文本理解而日益受关注。然而现有数据集多聚焦英文句级图像辅助抽取,缺乏视频驱动的细粒度视觉定位与多语言支持。为此,我们构建了名为M³D的多模态、多语言、多任务数据集:包含文档级文本与视频配对,支持英语和中文,涵盖实体识别、实体链抽取、关系抽取及视觉定位等任务,并引入传记主题拓展应用场景。为建立基准,我们提出分层多模态IE模型,通过去噪特征融合模块(DFFM)有效整合多模态信息;针对模态缺失问题,设计缺失模态构造模块(MMCM)。该模型在四个任务上于英中文数据集分别取得53.80%和53.77%的平均性能,为后续研究提供合理基线。进一步分析实验验证了各模块的有效性。

原文摘要 · Abstract (English)

Multimodal information extraction (IE) tasks have attracted increasing attention because many studies have shown that multimodal information benefits text information extraction. However, existing multimodal IE datasets mainly focus on sentence-level image-facilitated IE in English text, and pay little attention to video-based multimodal IE and fine-grained visual grounding. Therefore, in order to promote the development of multimodal IE, we constructed a multimodal multilingual multitask dataset, named M$^{3}$D, which has the following features: (1) It contains paired document-level text and video to enrich multimodal information; (2) It supports two widely-used languages, namely English and Chinese; (3) It includes more multimodal IE tasks such as entity recognition, entity chain extraction, relation extraction and visual grounding. In addition, our dataset introduces an unexplored theme, i.e., biography, enriching the domains of multimodal IE resources. To establish a benchmark for our dataset, we propose an innovative hierarchical multimodal IE model. This model effectively leverages and integrates multimodal information through a Denoised Feature Fusion Module (DFFM). Furthermore, in non-ideal scenarios, modal information is often incomplete. Thus, we designed a Missing Modality Construction Module (MMCM) to alleviate the issues caused by missing modalities. Our model achieved an average performance of 53.80% and 53.77% on four tasks in English and Chinese datasets, respectively, which set a reasonable standard for subsequent research. In addition, we conducted more analytical experiments to verify the effectiveness of our proposed module. We believe that our work can promote the development of the field of multimodal IE.

多模态信息抽取视频理解多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。