arXiv:2503.20988cs.CLcs.GR2025-03

用图结构+状态空间模型,高效理解图文数据并生成可解释摘要

Cross-Modal State-Space Graph Reasoning for Structured Summarization

  • 构建图文关系图,融合状态空间模型实现跨模态推理
  • 在多个基准上提升摘要质量,计算开销低于传统方法
  • 适合需要高效、可解释多模态摘要的场景

从大规模多模态数据中提取简洁、有意义的摘要对视频分析、医疗报告等应用至关重要。现有跨模态摘要方法常面临计算开销高和可解释性差的问题。本文提出一种交叉模态状态空间图推理(CSS-GR)框架,将状态空间模型与基于图的消息传递相结合,借鉴高效状态空间模型的设计思想。不同于依赖纯序列模型的方法,该框架构建图结构以捕捉文本与视觉流之间的跨模态及模态内关系,支持更全面的联合推理。实验表明,该方法在标准多模态摘要基准上显著提升了摘要质量与可解释性,同时保持了计算效率。我们还进行了详尽的消融研究,验证各组件的贡献。

原文摘要 · Abstract (English)

The ability to extract compact, meaningful summaries from large-scale and multimodal data is critical for numerous applications, ranging from video analytics to medical reports. Prior methods in cross-modal summarization have often suffered from high computational overheads and limited interpretability. In this paper, we propose a \textit{Cross-Modal State-Space Graph Reasoning} (\textbf{CSS-GR}) framework that incorporates a state-space model with graph-based message passing, inspired by prior work on efficient state-space models. Unlike existing approaches relying on purely sequential models, our method constructs a graph that captures inter- and intra-modal relationships, allowing more holistic reasoning over both textual and visual streams. We demonstrate that our approach significantly improves summarization quality and interpretability while maintaining computational efficiency, as validated on standard multimodal summarization benchmarks. We also provide a thorough ablation study to highlight the contributions of each component.

多模态摘要图神经网络状态空间模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。