arXiv:2510.07993cs.CLcs.AI2025-10被引 1

让论文图注既准确又符合作者写作风格。

Leveraging Author-Specific Context for Scientific Figure Caption Generation: 3rd SciCap Challenge

  • 分两阶段生成:先选上下文,再用作者风格微调。
  • 类别提示词使ROUGE-1召回率提升8.3%,精度损失仅2.8%。
  • 结合作者风格后,BLEU得分提高40%-48%,适合学术写作场景。

科学图表说明需兼具准确性与风格一致性。本文针对第三届SciCap挑战赛,提出一种领域专用的图注生成系统,融合图表相关文本上下文与作者特定写作风格,基于LaMP-Cap数据集。方法采用两阶段流程:第一阶段整合上下文过滤、类别特异性提示优化(通过DSPy的MIPROv2与SIMBA)及候选图注选择;第二阶段使用少量示例提示与作者画像进行风格精细化调整。实验表明,类别特异性提示优于零样本及通用优化方法,在保持精度损失仅-2.8%的前提下,使ROUGE-1召回率提升+8.3%,且BLEU-4下降-10.9%。引入作者画像风格微调后,BLEU得分提升40–48%,ROUGE得分提升25–27%。结果证明,结合上下文理解与作者风格适配,可生成既科学准确又风格忠实的图注。

原文摘要 · Abstract (English)

Scientific figure captions require both accuracy and stylistic consistency to convey visual information. Here, we present a domain-specific caption generation system for the 3rd SciCap Challenge that integrates figure-related textual context with author-specific writing styles using the LaMP-Cap dataset. Our approach uses a two-stage pipeline: Stage 1 combines context filtering, category-specific prompt optimization via DSPy's MIPROv2 and SIMBA, and caption candidate selection; Stage 2 applies few-shot prompting with profile figures for stylistic refinement. Our experiments demonstrate that category-specific prompts outperform both zero-shot and general optimized approaches, improving ROUGE-1 recall by +8.3\% while limiting precision loss to -2.8\% and BLEU-4 reduction to -10.9\%. Profile-informed stylistic refinement yields 40--48\% gains in BLEU scores and 25--27\% in ROUGE. Overall, our system demonstrates that combining contextual understanding with author-specific stylistic adaptation can generate captions that are both scientifically accurate and stylistically faithful to the source paper.

图注生成风格适配SciCap两阶段模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。