让AI生成更符合作者风格的图文说明,用多图上下文提升准确性。
LaMP-Cap: Personalized Figure Caption Generation With Multimodal Figure Profiles
- 基于多模态图表上下文生成个性化描述
- 使用图表关联图像比文字段落更有效提升质量
- 适合需要精准表达的科研写作者和学术出版
图表标题对帮助读者理解与记忆图表核心信息至关重要。尽管已有多种模型用于自动生成标题,作者仍需大量修改通用AI生成内容以匹配自身写作风格与领域习惯,凸显个性化需求。现有语言模型的个性化技术多限于纯文本场景,极少涉及输入与个人资料均为多模态的情况。本文提出LaMP-Cap数据集,支持基于多模态图表资料的个性化标题生成。每个目标图表不仅包含图像,还提供同文档内最多三张相关图表——每张含图像、标题及提及该图的段落——作为上下文资料。四项大模型实验表明,引入这些资料可显著提升生成标题与原始作者撰写版本的相似度。消融实验证明,资料中的图像比文字段落更具价值,凸显多模态资料相比纯文本的优势。
原文摘要 · Abstract (English)
Figure captions are crucial for helping readers understand and remember a figure's key message. Many models have been developed to generate these captions, helping authors compose better quality captions more easily. Yet, authors almost always need to revise generic AI-generated captions to match their writing style and the domain's style, highlighting the need for personalization. Despite language models' personalization (LaMP) advances, these technologies often focus on text-only settings and rarely address scenarios where both inputs and profiles are multimodal. This paper introduces LaMP-Cap, a dataset for personalized figure caption generation with multimodal figure profiles. For each target figure, LaMP-Cap provides not only the needed inputs, such as figure images, but also up to three other figures from the same document--each with its image, caption, and figure-mentioning paragraphs--as a profile to characterize the context. Experiments with four LLMs show that using profile information consistently helps generate captions closer to the original author-written ones. Ablation studies reveal that images in the profile are more helpful than figure-mentioning paragraphs, highlighting the advantage of using multimodal profiles over text-only ones.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。