构建首个跨科学媒体的对齐数据集,实现论文、幻灯片、视频间细粒度对应。
Unifying Scientific Communication: Fine-Grained Correspondence Across Scientific Media

- 建立论文、幻灯片、视频等多模态科学内容的统一数据集
- 发现视觉-语言模型在细粒度对齐上表现有限,符号内容易形成孤立聚类
- 适合研究多模态知识表示与科学传播的学者参考
科学知识的传播日益呈现多模态特征,涵盖文本、图像和语音等形式,如研究论文、演示幻灯片和录制演讲。这些不同形式共同传达研究的推理过程、结果与洞见,提供互补视角以深化理解。然而,尽管目标一致,各形式间缺乏结构化关联。由于缺乏显式链接,难以追踪概念、图表与解释之间的对应关系,限制了对研究内容的统一探索与分析。为此,我们提出首个整合同一研究作品的论文、演示视频、解说视频和幻灯片的多模态会议数据集(MCD),并评估多种基于嵌入和视觉-语言模型在发现跨格式细粒度对应关系上的能力,建立了该任务的首个系统性基准。结果表明,视觉-语言模型具有鲁棒性但难以实现细粒度对齐;基于嵌入的方法能较好捕捉文本-图像对应,但公式与符号内容在嵌入空间中形成独立聚类。这些发现揭示了现有方法的优势与局限,并指明未来多模态科学理解的研究方向。为保障可复现性,相关资源已发布于 https://github.com/meghamariamkm2002/MCD。
原文摘要 · Abstract (English)
The communication of scientific knowledge has become increasingly multimodal, spanning text, visuals, and speech through materials such as research papers, slides, and recorded presentations. These different representations collectively convey a study's reasoning, results, and insights, offering complementary perspectives that enrich understanding. However, despite their shared purpose, such materials are rarely connected in a structured way. The absence of explicit links across formats makes it difficult to trace how concepts, visuals, and explanations correspond, limiting unified exploration and analysis of research content. To address this gap, we introduce the Multimodal Conference Dataset (MCD), the first benchmark that integrates research papers, presentation videos, explanatory videos, and slides from the same works. We evaluate a range of embedding-based and vision-language models to assess their ability to discover fine-grained cross-format correspondences, establishing the first systematic benchmark for this task. Our results show that vision-language models are robust but struggle with fine-grained alignment, while embedding-based models capture text-visual correspondences well but equations and symbolic content form distinct clusters in the embedding space. These findings highlight both the strengths and limitations of current approaches and point to key directions for future research in multimodal scientific understanding. To ensure reproducibility, we release the resources for MCD at https://github.com/meghamariamkm2002/MCD
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。