提出可控多模态融合框架OTCR,提升文档信息抽取精度
OTCR: Optimal Transmission, Compression and Representation for Multimodal Information Extraction
- 用最优传输实现文本与视觉稀疏对齐,动态控制视觉信息注入
- 在FUNSD和XFUND(ZH)上分别达91.95%和94.20%的抽取准确率
- 适合需要可解释、低冗余多模态融合的文档智能任务
多模态信息抽取(MIE)需融合图文富集文档中的文本与视觉线索。现有方法多隐含假设模态等价或采用统一融合范式,导致无关冗余信息混入,限制泛化能力。本文从任务导向视角出发:文本主导,视觉选择性支持。提出OTCR两阶段框架:首先通过跨模态最优传输(OT)生成文本词元与视觉块间的稀疏概率对齐,并以上下文感知门控控制视觉注入;其次利用变分信息瓶颈(VIB)压缩融合特征,过滤任务无关噪声,生成紧凑且任务自适应的表示。在FUNSD数据集上达到91.95%的序列级精确率(SER)与91.13%的召回率(RE),在XFUND(ZH)上达91.09% SER与94.20% RE,性能具有竞争力。特征层面分析证实模态冗余降低、任务信号增强。本工作为文档AI中的可控多模态融合提供了可解释、信息论驱动的新范式。
原文摘要 · Abstract (English)
Multimodal Information Extraction (MIE) requires fusing text and visual cues from visually rich documents. While recent methods have advanced multimodal representation learning, most implicitly assume modality equivalence or treat modalities in a largely uniform manner, still relying on generic fusion paradigms. This often results in indiscriminate incorporation of multimodal signals and insufficient control over task-irrelevant redundancy, which may in turn limit generalization. We revisit MIE from a task-centric view: text should dominate, vision should selectively support. We present OTCR, a two-stage framework. First, Cross-modal Optimal Transport (OT) yields sparse, probabilistic alignments between text tokens and visual patches, with a context-aware gate controlling visual injection. Second, a Variational Information Bottleneck (VIB) compresses fused features, filtering task-irrelevant noise to produce compact, task-adaptive representations. On FUNSD, OTCR achieves 91.95% SER and 91.13% RE, while on XFUND (ZH), it reaches 91.09% SER and 94.20% RE, demonstrating competitive performance across datasets. Feature-level analyses further confirm reduced modality redundancy and strengthened task signals. This work offers an interpretable, information-theoretic paradigm for controllable multimodal fusion in document AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。