构建首个数据视频自动生成基准,评估模型讲故事与动画同步能力
DATAREEL: Automated Data-Driven Video Story Generation with Animations
- 设计分步生成流程:规划-编码-验证,提升生成质量
- 开放模型执行失败率高达39.8%,普遍存在画面静止与字幕不同步
- 适合关注数据可视化生成、智能内容创作的研究者使用
数据视频通过动画可视化与同步旁白传递量化信息,广泛应用于新闻、教育和公共传播。自动构建此类视频需决定讲述故事、设计有效可视化并生成与旁白同步的可执行动画。尽管视觉语言模型(VLMs)进展迅速,但其从高层次传播意图生成数据视频的能力仍不明确,因缺乏标准化基准。我们提出DATAREEL,一个包含328个真实世界数据视频片段的基准。给定数据表、传播意图、目标时长和风格参考图,模型需生成可执行动画代码及同步字幕,我们进行渲染与评估。测试八种专有与开源VLM发现显著能力差距:开源模型执行失败率最高达39.8%,而专有模型表现更优。但即使最优模型仍常生成静态图表、字幕-动画不同步、布局不稳定及偏离参考风格。我们进一步提出一种强效代理基线,通过规划、编码、验证分解流程,在人工与自动评估中均优于直接提示。该任务仍未解决;我们已将DATAREEL发布于https://github.com/vis-nlp/DataReel以支持后续研究。
原文摘要 · Abstract (English)
Data videos combine animated visualizations with synchronized narration to communicate quantitative information and are widely used in journalism, education, and public communication. Automatically generating them requires deciding what story to tell, designing effective visualizations, and producing executable animations synchronized with narration. Despite rapid progress in vision-language models (VLMs), it remains unclear how well they can perform this task from a high-level communicative intent, largely because no standardized benchmark exists. We introduce DATAREEL, a benchmark for automated data-driven video story generation containing 328 real-world data reels. Given a data table, a communicative intent, a target duration, and a style reference image, a model must generate executable animation code with synchronized subtitles, which we render and evaluate. Evaluating eight proprietary and open-weight VLMs reveals a substantial capability gap: open-weight models exhibit execution failure rates of up to 39.8%, whereas proprietary models achieve stronger overall performance. Yet even the best-performing models frequently generate static charts, subtitle-animation desynchronization, unstable layouts, and poor adherence to the reference style. We further introduce a strong agentic baseline that decomposes generation into planning, coding, and verification, consistently outperforming direct prompting in both human and automatic evaluations. The task remains far from solved; we release DATAREEL at https://github.com/vis-nlp/DataReel to support future work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。