用自然语言生成符合场景几何的自动动画路径
Generative Animations: A Multi-Model Pipeline for Prompt-Driven Motion Synthesis
- 结合LLM与SAM模型,将文本提示转为视觉对齐的运动轨迹
- 支持轮廓跟随、环绕动画和透视对齐运动三种场景
- 适合需要快速生成复杂动画的设计师或内容创作者
动画能将数字文档变为沉浸式体验,但自定义运动路径仍需手动选择预设、绘制贝塞尔点并配置时间属性。我们提出生成式动画系统,可将自然语言提示直接转化为可投入生产的动画。通过串联大语言模型(LLM)进行语义解析,以及使用分割一切模型(SAM)实现视觉定位,该流程能自动生成符合场景几何、处理深度遮挡并遵循三维透视变换的运动路径。我们在三个应用场景中验证系统:轮廓跟踪轨迹、具备层级顺序感知的环绕动画,以及在变形物体上的透视对齐运动。
原文摘要 · Abstract (English)
Animation elevates digital documents into immersive experiences, yet creating custom motion paths remains cumbersome, requiring designers to manually select presets, plot Bézier points, and configure timing properties. We introduce Generative Animations, a system that transforms natural language prompts into production-ready animations. By chaining Large Language Models (LLMs) for semantic parsing with the Segment Anything Model (SAM) for visual grounding, our pipeline automatically generates motion paths that respect scene geometry, handle depth-based occlusions, and honor 3D perspective transforms. We demonstrate the system through three use cases: contour-following trajectories, orbital animations with z-order awareness, and perspective-aligned motion on transformed objects.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。