构建科学海报摘要基准,推动多模态模型理解复杂图文内容
PosterSum: A Multimodal Benchmark for Scientific Poster Summarization
- 设计分层摘要方法,先分割后生成提升理解精度
- 在16,305张海报上测试,模型摘要准确率仍有显著提升空间
- 适合研究多模态生成、学术智能系统的开发者参考
从多模态文档中生成准确简洁的文本摘要极具挑战性,尤其面对科学海报这类视觉结构复杂的材料。我们提出PosterSum,一个新型基准数据集,旨在推进视觉-语言模型对科学海报的理解与摘要能力。该数据集包含16,305张会议海报及其对应的论文摘要作为参考。每张海报以图像形式提供,具有复杂布局、密集文本区域、表格和图表等多样视觉挑战。我们在PosterSum上评估了当前先进的多模态大语言模型(MLLMs),发现其在准确解析和摘要科学海报方面表现不佳。为此,我们提出一种分段-摘要的层级方法,优于现有MLLMs,在自动评价指标上实现ROUGE-L提升3.14%。这将为未来海报摘要研究提供起点。
原文摘要 · Abstract (English)
Generating accurate and concise textual summaries from multimodal documents is challenging, especially when dealing with visually complex content like scientific posters. We introduce PosterSum, a novel benchmark to advance the development of vision-language models that can understand and summarize scientific posters into research paper abstracts. Our dataset contains 16,305 conference posters paired with their corresponding abstracts as summaries. Each poster is provided in image format and presents diverse visual understanding challenges, such as complex layouts, dense text regions, tables, and figures. We benchmark state-of-the-art Multimodal Large Language Models (MLLMs) on PosterSum and demonstrate that they struggle to accurately interpret and summarize scientific posters. We propose Segment & Summarize, a hierarchical method that outperforms current MLLMs on automated metrics, achieving a 3.14% gain in ROUGE-L. This will serve as a starting point for future research on poster summarization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。