arXiv:2608.02892cs.CV2026-08

针对科研实验场景的复杂关系,提出新模型与数据集。

Modeling Scientific Experiment Scenes: Dataset and Model

论文配图:Modeling Scientific Experiment Scenes: Dataset and Model
图 1 · 摘自论文原文
  • 融合视觉与文本信息增强物体表征,利用几何线索提升关系推理。
  • 在PhysScene数据集上实现优于现有方法的开放词汇场景图生成效果。
  • 适合关注科学图像理解、多模态学习的研究者使用。

场景图生成(SGG)是结构化视觉理解的基础,但现有基准主要聚焦日常生活图像,忽视了包含专业仪器、任务特定语义及密集细粒度物理关系的科研实验场景。基于我们此前提出的物理实验场景SGG数据集PhysScene,我们识别出两大关键挑战:显著的长尾关系谓词分布和较大的视觉-文本语义鸿沟。为此,我们提出跨模态双路径生成器(CM-DPG),一种鲁棒的开放词汇SGG模型。该模型通过联合视觉-文本编码增强物体语义表示,并利用互补的视觉与几何线索改进关系推理。同时引入关系感知预训练、基于描述的伪监督和自适应加权机制,以平衡头尾谓词的学习。在PhysScene和VG150上的大量实验表明,CM-DPG在多种评估设置下表现优异,消融实验证明了各组件的有效性。数据集与代码已公开于https://github.com/ZMH-SDUST/CM-DPG。

原文摘要 · Abstract (English)

Scene Graph Generation (SGG) is fundamental to structured visual understanding, yet existing benchmarks focus mainly on daily-life images and overlook scientific experiment scenes with specialized instruments, task-specific experimental semantics, and dense, fine-grained physical relations. Building upon PhysScene, our previously introduced SGG dataset for physics experiment scenes, we further identify two key challenges that such scientific environments pose to existing SGG models: a pronounced long-tail relational predicate distribution and a substantial visual-textual semantic gap. To address these challenges, we propose the Cross-Modal Dual-Path Generator (CM-DPG), a model for robust open-vocabulary SGG. The model enhances object-level semantic representations through joint visual-textual encoding and improves relational reasoning using complementary visual and geometric cues. We also incorporate relation-aware pre-training, caption-derived pseudo-supervision, and adaptive weighting to support balanced learning across head and tail predicates. Extensive experiments on PhysScene and VG150 show that CM-DPG achieves competitive performance across multiple evaluation settings, with ablation studies validating the contribution of each component. The dataset and code are publicly available at https://github.com/ZMH-SDUST/CM-DPG.

场景图生成科研图像多模态学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。