用多模态大模型消除3D/4D生成中的几何错乱和时间抖动。
Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation

- 通过生成-检测-修正范式,用大语言模型识别多视角多帧的不一致。
- 在不重训练的前提下,实现跨视角和时序的几何与外观一致性优化。
- 适合需要高一致性3D/4D内容生成的研究者和开发者使用。
尽管近期3D生成技术已实现出色的视觉合成,但现有方法通常依赖2D扩散监督,缺乏显式的几何一致性机制,导致空间幻觉(如结构重复、对齐错误)。这些问题在4D生成中更为严重,因需保持视角间和时间演变中的一致性,引发抖动、身份闪烁和结构漂移。本文提出统一且模型无关的Hallo4D框架,用于缓解时空幻觉。Hallo4D采用生成-检测-修正范式,利用大得多模态语言模型(LMMs)从多视角和多帧渲染中识别并总结时空不一致。这些洞察指导基于共识的图像空间一致性优化,其中LMM-based选择器通过多模型投票评估候选修正方案,无需重新训练或修改架构。为进一步提升时序一致性和优化效率,系统引入运动感知关键帧采样、LMM引导初始化和外观对齐。此外,还设计曝光感知优化与可见性剪枝以增强复杂视角下的鲁棒性。大量实验表明,Hallo4D在多种3D和4D生成场景中持续优于强基线,提供可扩展、通用的一致性感知内容生成方案。
原文摘要 · Abstract (English)
While recent advances in 3D generation have enabled impressive visual synthesis, existing methods often rely on 2D diffusion supervision without explicit mechanisms for geometric consistency, leading to spatial hallucinations such as duplicated structures and misaligned geometry. These issues become more severe in 4D generation, where maintaining consistency across viewpoints and temporal evolution introduces additional challenges, including jitter, identity flicker, and structural drift. We present \textbf{Hallo4D}, a unified and model-agnostic framework for mitigating spatiotemporal hallucinations in 3D and 4D content generation. Hallo4D introduces a generation-detection-correction paradigm that leverages large multimodal language models (LMMs) to identify and summarize spatial and temporal inconsistencies from multi-view and multi-frame renderings. These insights guide a consensus-driven image-space consistency optimization, where an LMM-based selector evaluates candidate corrections through multi-model voting, without requiring retraining or architectural modifications. To further improve temporal consistency and optimization efficiency, Hallo4D incorporates motion-aware keyframe sampling, LMM-guided initialization, and appearance alignment. We additionally introduce exposure-aware optimization and visibility pruning to enhance robustness under challenging viewpoints. Extensive experiments demonstrate that Hallo4D consistently outperforms strong baselines across diverse 3D and 4D generation settings, providing a scalable and generalizable solution for consistency-aware content generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。