让文字和图片更好融合,生成更准确的多模态摘要。
Towards Visually Grounded Multimodal Summarization via Cross-Modal Transformer and Gated Attention

- 深度对齐视觉与语言模型,实现分层融合。
- 通过软标签蒸馏选关键图像,提升代表性。
- 适合需要图文协同理解的任务场景。
多模态摘要需联合理解文本与视觉输入以生成简洁、语义连贯的摘要。现有方法常将浅层视觉特征注入深层语言模型,导致表征不匹配和跨模态关联弱。我们提出统一框架,同步完成文本摘要与代表性图像选择。系统SPeCTrA-Sum(Sampler Perceiver with Cross-modal Transformer and gated Attention for Summarization)引入两项创新:首先,深度视觉处理器(DVP)使视觉编码器与语言模型在对应层次对齐,实现分层、逐层融合,保障语义一致性;其次,轻量级视觉相关性预测器(VRP)通过从确定性点过程(DPP)教师模型中蒸馏软标签,筛选出显著且多样化的图像。SPeCTrA-Sum采用多目标损失训练,包含自回归摘要、跨模态对齐与基于DPP的蒸馏。实验表明,该系统生成的摘要更准确、更具视觉依据,且所选图像更具代表性,验证了深度感知融合与原则化图像选择在多模态摘要中的优势。
原文摘要 · Abstract (English)
Multimodal summarization requires models to jointly understand textual and visual inputs to generate concise, semantically coherent summaries. Existing methods often inject shallow visual features into deep language models, leading to representational mismatches and weak cross-modal grounding. We propose a unified framework that jointly performs text summarization and representative image selection. Our system, SPeCTrA-Sum (Sampler Perceiver with Cross-modal Transformer and gated Attention for Summarization), introduces two key innovations. First, a Deep Visual Processor (DVP) aligns the visual encoder with the language model at corresponding depths, enabling hierarchical, layer-wise fusion that preserves semantic consistency. Second, a lightweight Visual Relevance Predictor (VRP) selects salient and diverse images by distilling soft labels from a Determinantal Point Processes (DPP) teacher. SPeCTrA-Sum is trained using a multi-objective loss that combines autoregressive summarization, cross-modal alignment, and DPP-based distillation. Experiments show that our system produces more accurate, visually grounded summaries and selects more representative images, demonstrating the benefits of depth-aware fusion and principled image selection for multimodal summarization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。