首个幻灯片动画数据集,让AI理解并生成动态幻灯片效果。
Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models
- 构建12000组图文动画三元组,覆盖所有PowerPoint内置动画。
- 微调模型在多项指标上超越GPT-4.1和Gemini-2.5-Pro,BLEU-4提升60%。
- 提出CODA评估指标,量化动作覆盖、时序顺序与细节保真度。
幻灯片动画(如淡入、飞入、擦除)对观众吸引力、信息传达效率和视觉表现力至关重要。然而,当前多数AI幻灯片生成工具缺乏原生动画支持,现有视觉语言模型因缺少公开数据集和有限的时间推理能力而难以处理动画任务。为此,我们发布了首个幻灯片动画建模公开数据集:包含12,000个自然语言描述、动画JSON文件和渲染视频的三元组,涵盖所有内置PowerPoint动画效果。基于该数据集,我们使用低秩适配(LoRA)微调Qwen-2.5-VL-7B模型,在BLEU-4、ROUGE-L、SPICE及自研的覆盖率-时序-细节评估(CODA)指标上均优于GPT-4.1和Gemini-2.5-Pro。在人工构建的测试集上,模型使BLEU-4提升约60%,ROUGE-L提升30%,且在CODA细节维度表现显著增强。结果表明,低秩适配可有效实现可靠的时间推理与泛化能力。整体而言,本研究提供了数据、模型与评估基准,为未来基于VLM的动态幻灯片生成研究奠定基础。
原文摘要 · Abstract (English)
Slide animations, such as fade-in, fly-in, and wipe, are critical for audience engagement, efficient information delivery, and vivid visual expression. However, most AI-driven slide-generation tools still lack native animation support, and existing vision-language models (VLMs) struggle with animation tasks due to the absence of public datasets and limited temporal-reasoning capabilities. To address this gap, we release the first public dataset for slide-animation modeling: 12,000 triplets of natural-language descriptions, animation JSON files, and rendered videos, collectively covering every built-in PowerPoint effect. Using this resource, we fine-tune Qwen-2.5-VL-7B with Low-Rank Adaptation (LoRA) and achieve consistent improvements over GPT-4.1 and Gemini-2.5-Pro in BLEU-4, ROUGE-L, SPICE, and our Coverage-Order-Detail Assessment (CODA) metric, which evaluates action coverage, temporal order, and detail fidelity. On a manually created test set of slides, the LoRA model increases BLEU-4 by around 60%, ROUGE-L by 30%, and shows significant improvements in CODA-detail. This demonstrates that low-rank adaptation enables reliable temporal reasoning and generalization beyond synthetic data. Overall, our dataset, LoRA-enhanced model, and CODA metric provide a rigorous benchmark and foundation for future research on VLM-based dynamic slide generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。