用单视角视频让2D模型玩转4D场景交互,无需3D数据
Feature4X: Bridging Any Monocular Video to 4D Agentic AI with Versatile Gaussian Feature Fields

- 通过高斯点云将2D模型特征蒸馏到4D空间,实现跨时空统一表征
- 支持任意视角分割、动态场景编辑和全时序视觉问答,效果超越现有方法
- 适合想用普通视频做智能交互的开发者和研究者
近年来,2D与多模态模型在大规模数据训练下取得了显著进展。然而,将其能力扩展至复杂3D/4D场景的自由交互与高层语义操作仍面临挑战,主要受限于高质量标注的3D/4D或多视角数据集稀缺。本文提出Feature4X,一个通用框架,仅需单视角视频输入(如用户生成内容),即可将任意2D视觉基础模型功能延伸至4D领域。'X'代表其灵活性,支持通过模型自适应的4D特征场蒸馏实现各类任务。核心是动态优化策略,将多种模型能力统一为单一表示。据我们所知,Feature4X是首个使用高斯点阵将视频基础模型(如SAM2、InternVideo2)特征显式蒸馏至4D特征场的方法。实验展示了新颖的任意视角分割、几何与外观场景编辑及全时步自由形式视觉问答,由大语言模型在反馈环中驱动。这些进展拓展了代理型AI应用边界,为可扩展、时空感知的沉浸式动态4D场景交互提供了基础。
原文摘要 · Abstract (English)
Recent advancements in 2D and multimodal models have achieved remarkable success by leveraging large-scale training on extensive datasets. However, extending these achievements to enable free-form interactions and high-level semantic operations with complex 3D/4D scenes remains challenging. This difficulty stems from the limited availability of large-scale, annotated 3D/4D or multi-view datasets, which are crucial for generalizable vision and language tasks such as open-vocabulary and prompt-based segmentation, language-guided editing, and visual question answering (VQA). In this paper, we introduce Feature4X, a universal framework designed to extend any functionality from 2D vision foundation model into the 4D realm, using only monocular video input, which is widely available from user-generated content. The "X" in Feature4X represents its versatility, enabling any task through adaptable, model-conditioned 4D feature field distillation. At the core of our framework is a dynamic optimization strategy that unifies multiple model capabilities into a single representation. Additionally, to the best of our knowledge, Feature4X is the first method to distill and lift the features of video foundation models (e.g., SAM2, InternVideo2) into an explicit 4D feature field using Gaussian Splatting. Our experiments showcase novel view segment anything, geometric and appearance scene editing, and free-form VQA across all time steps, empowered by LLMs in feedback loops. These advancements broaden the scope of agentic AI applications by providing a foundation for scalable, contextually and spatiotemporally aware systems capable of immersive dynamic 4D scene interaction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。