让无字或短字科学图版自动生成文字说明,提升论文图表可用性。
FigEx2: Visual-Conditioned Panel Detection and Captioning for Scientific Compound Figures
- 用视觉信息直接定位图版并生成描述,无需依赖原始文本。
- 检测准确率[email protected]:0.95达0.728,生成质量优于Qwen3-VL-8B。
- 适用于跨学科科学图像,零样本迁移能力强,适合科研数据预训练。
科学复合图将多个带标签的图版整合为单张图像。在对346,567张复合图的PMC规模爬取中,16.3%无文字说明,1.8%说明不足十个词,导致被现有拆解流程丢弃。本文提出FigEx2,一种视觉条件化框架,直接从图像中定位图版并生成图版级描述,将原本不可用的图像转化为可对齐的图文对,用于下游预训练与检索。为缓解开放式描述中的语言差异,引入噪声感知门控融合模块,自适应调节描述特征对检测查询空间的控制;采用分阶段SFT+RL策略,结合CLIP对齐与BERTScore语义奖励。为支持高质量监督,构建BioSci-Fig-Cap基准,涵盖生物、物理、化学跨领域测试集。FigEx2在检测任务中达到0.728 [email protected]:0.95,METEOR优于Qwen3-VL-8B 0.44,BERTScore高0.22,并可在无微调情况下零样本迁移至分布外科学领域。
原文摘要 · Abstract (English)
Scientific compound figures combine multiple labeled panels into a single image. However, in a PMC-scale crawl of 346,567 compound figures, 16.3% have no caption and 1.8% only have captions shorter than ten words, causing them to be discarded by existing caption-decomposition pipelines. We propose FigEx2, a visual-conditioned framework that localizes panels and generates panel-wise captions directly from the image, converting otherwise unusable figures into aligned panel-text pairs for downstream pretraining and retrieval. To mitigate linguistic variance in open-ended captioning, we introduce a noise-aware gated fusion module that adaptively controls how caption features condition the detection query space, and employ a staged SFT+RL strategy with CLIP-based alignment and BERTScore-based semantic rewards. To support high-quality supervision, we curate BioSci-Fig-Cap, a refined benchmark for panel-level grounding, alongside cross-disciplinary test suites in physics and chemistry. FigEx2 achieves 0.728 [email protected]:0.95 for detection, outperforms Qwen3-VL-8B by 0.44 in METEOR and 0.22 in BERTScore, and transfers zero-shot to out-of-distribution scientific domains without fine-tuning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。