解决视频生成4D网格时的复杂拓扑变化难题
Helix4D: Complex 4D Mesh Generation

- 用滑动窗口跨帧注意力融合时间信息,保持初始帧质量
- 4D时间编码复用空间位置编码的低频带,零参数扩展
- 在动作与复杂动态数据集上生成高质量动态网格
当前视频到4D网格的方法在处理复杂拓扑变化、透明材质、细长结构和内部表面时表现不佳。我们提出Helix4D,一个基于Trellis2表达能力的动态网格生成框架,将其从图像到3D扩展为视频条件下的4D生成。核心问题在于:(a) 如何让Trellis2的帧内注意力在跨帧间共享信息,同时保留其对稀有情况(如透明物体、内部表面)的预训练质量;(b) 如何向纯3D位置编码注入时间信息而不破坏预训练能力。针对(a),采用滑动窗口跨帧注意力,并以首帧作为锚点——首帧由基础Trellis2模型生成并注入模型,通过跨帧注意力继承其高质量表现。针对(b),引入4D时间编码,将冗余的低频空间RoPE频带重用于时间维度,实现3D编码向4D无参扩展。大量实验表明,Helix4D在ActionBench和自建复杂动态数据集上均能生成高质量动态网格。
原文摘要 · Abstract (English)
Current video-to-4D methods struggle with complex topology changes, transparent materials, thin structures, and inner surfaces. We present Helix4D, a dynamic mesh generation framework by inheriting the expressive representation of Trellis2, adapting it from image-to-3D to video-conditioned 4D generation. Our design arises from two key questions: (a) how to enable Trellis2's frame-local attention to share information across frames while preserving its pretrained quality on rare cases such as transparent objects and inner surfaces, and (b) how to inject temporal information into a purely 3D positional encoding without breaking pretrained capabilities. We address (a) with a sliding-window cross-frame attention and anchor on the first frame. The first frame is generated by the base Trellis2 model and injected into our model, letting it inherit Trellis2's quality in rare cases through cross-frame attention. We address (b) with a 4D temporal encoding that repurposes redundant low-frequency spatial RoPE bands for time, extending the encoding from 3D with no additional parameters. Extensive experiments show the effectiveness of Helix4D for high-quality dynamic mesh generation on ActionBench and our own challenging complex dynamics set.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。