arXiv:2606.18780cs.CVcs.CL2026-06中稿 · IEEE Transactions …

用语义锚点统一生成多模态低资源信息抽取数据

SAMA: Semantic Anchor-aligned Augmentation for Unified Low-Resource Multimodal Information Extraction

论文配图:SAMA: Semantic Anchor-aligned Augmentation for Unified Low-Resource Multimodal Information Extraction
图 1 · 摘自论文原文
  • 以真实标签构建语义锚点,指导生成符合任务需求的合成数据
  • 在多个基准数据集上显著提升低资源场景下的抽取性能
  • 适合需要高效数据增强的多模态信息抽取研究者使用

多模态信息抽取(MIE)涵盖多模态命名实体识别(MNER)、关系抽取(MRE)和事件抽取(MEE),对理解多媒体内容至关重要,但受限于严重数据稀缺。尽管数据增强是有效方案,现有方法因跨模态对齐粗糙、任务专用设计碎片化而难以利用共享语义知识。为此,本文提出统一的语义锚点对齐多模态增强框架SAMA。SAMA从真实标签构建结构化语义锚点,引导协作式多专家多模态大模型(CME-MLLM)生成高质量文本样本;该模型融合通用适配器与任务特定适配器,实现多样性与约束一致性平衡。图像生成采用锚点保持扩散机制,通过锚点加权提示与潜在条件保持关键语义锚点,同时丰富视觉上下文。为避免人工验证,SAMA引入双约束过滤模块,基于跨模态一致性和锚点保真度筛选合成样本。在多个主流MNER、MRE和MEE数据集上的实验表明,SAMA在全监督与低资源设置下均优于现有最佳增强基线,证明其通用性、鲁棒性与有效性。

原文摘要 · Abstract (English)

Multimodal Information Extraction (MIE)-covering tasks such as Multimodal Named Entity Recognition (MNER), Relation Extraction (MRE), and Event Extraction (MEE)-is essential for understanding multimedia content but remains constrained by severe data scarcity. Although data augmentation is a promising remedy, existing approaches are impeded by coarse cross-modal alignment and fragmented, task-specific designs that fail to exploit shared semantic knowledge. To overcome these limitations, we introduce Semantic Anchor-aligned Multimodal Augmentation (SAMA), a unified framework for generating high-fidelity, task-aware synthetic data. SAMA constructs structured semantic anchors from ground-truth labels to guide a Collaborative Multi-Experts Multimodal Large Language Model (CME-MLLM), which integrates a Universal Adapter for shared semantics with Task-Specific Adapters to produce diverse yet constraint-compliant textual samples. For image synthesis, SAMA employs an Anchor-Preserving Diffusion mechanism that uses anchor-weighted prompts and latent conditioning to maintain critical semantic anchors while diversifying visual contexts. To eliminate the need for manual verification, SAMA further introduces a Dual-Constraint Filtering module that selects synthetic samples based on both cross-modal consistency and anchor fidelity. Extensive experiments across benchmark datasets for MNER, MRE, and MEE demonstrate that SAMA consistently outperforms state-of-the-art augmentation baselines under both fully supervised and low-resource settings, underscoring its versatility, robustness, and effectiveness.

多模态数据增强低资源信息抽取

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。