arXiv:2607.28006cs.AI2026-07

解决多模态长文档摘要中的关键信息遗漏和跨模态幻觉问题。

MMLDSum-LLM: Multimodal Long-Document Summarization with Visual-Alignment and Keyword-Aware

论文配图:MMLDSum-LLM: Multimodal Long-Document Summarization with Visual-Alignment and Keyword-Aware
图 1 · 摘自论文原文
  • 分两阶段训练,融合视觉对齐与关键词感知损失。
  • 在多领域长文档上实现更高关键词覆盖率和跨模态一致性。
  • 适合需要精准提取复杂多模态知识的科研与工程场景。

多模态长文档是专业领域知识的核心载体,关键证据常分散于不同段落与模态中,易导致多模态大模型摘要时出现关键信息遗漏与跨模态幻觉。根源在于长程依赖建模中的注意力漂移及模态间对齐缺失。为此,我们构建了MMLDSum-Bench——一个覆盖多领域、多上下文长度与视觉-文本模态分布的高质量基准数据集。进一步提出MMLDSum-LLM,一种可复现的两阶段训练框架:先通过监督微调结合视觉对齐加权损失与关键词感知加权损失,再采用GRPO优化多目标奖励(关键词覆盖率、图文对齐、ROUGE、长度控制)。在统一评估协议下,对比领先闭源与开源多模态模型,实验表明该方法显著提升关键信息覆盖率与跨模态一致性,涵盖LLM-as-a-judge评分、原子声明精确率/召回率、图像-文本对齐(ITA)与ROUGE等指标。

原文摘要 · Abstract (English)

Multimodal long documents are core carriers of professional knowledge, where critical evidence is sparsely distributed across paragraphs and modalities. This easily causes key information omission and cross-modal hallucinations in summarization by multimodal LLMs. These issues stem from attention drift in long-range dependency modeling and gaps in inter-modal alignment. To address this, we introduce MMLDSum-Bench, a high-quality benchmark for multimodal long-document summarization, covering multiple domains, context-length scales, and visual-textual modality distributions. We further propose MMLDSum-LLM, a reproducible two-stage training framework that combines supervised fine-tuning with visual-alignment weighted loss and keyword-aware weighted loss, followed by GRPO with a multi-objective reward (keyword coverage, image-text alignment, ROUGE, and length control). Extensive experiments on MMLDSum-Bench, comparing against leading closed-source and open-source multimodal models under a unified evaluation protocol - including LLM-as-a-judge scoring, atomic-claim precision/recall, image-text alignment (ITA), and ROUGE - demonstrate that our approach significantly improves key-information coverage and cross-modal consistency.

多模态摘要长文档视觉对齐关键词感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。