主动选择最有信息量的视角,提升3D高斯点云的语义与动态建模效果。
Next Best View Selections for Semantic and Dynamic 3D Gaussian Splatting
- 基于Fisher信息设计主动学习策略,量化视角对语义和变形参数的贡献。
- 在多相机数据上选帧,显著提升渲染质量与分割精度,优于随机与不确定度基线。
- 适合需要高效采集动态场景信息的机器人、自动驾驶等应用。
理解语义与动态特性对具身智能体完成各类任务至关重要,这两类任务的数据冗余远高于静态场景理解。本文将视角选择问题建模为主动学习问题,目标是优先选择能为模型训练带来最大信息增益的帧。为此,提出一种基于Fisher信息的主动学习算法,量化候选视角对语义高斯参数与形变网络的信息价值。该方法可联合处理语义推理与动态场景建模,提供比启发式或随机策略更合理的解决方案。我们在大规模静态图像与动态视频数据集上,从多相机设置中选择有信息量的帧进行评估。实验表明,本方法在渲染质量和语义分割性能上均持续优于基于随机选择与不确定性启发式的基线方法。
原文摘要 · Abstract (English)
Understanding semantics and dynamics has been crucial for embodied agents in various tasks. Both tasks have much more data redundancy than the static scene understanding task. We formulate the view selection problem as an active learning problem, where the goal is to prioritize frames that provide the greatest information gain for model training. To this end, we propose an active learning algorithm with Fisher Information that quantifies the informativeness of candidate views with respect to both semantic Gaussian parameters and deformation networks. This formulation allows our method to jointly handle semantic reasoning and dynamic scene modeling, providing a principled alternative to heuristic or random strategies. We evaluate our method on large-scale static images and dynamic video datasets by selecting informative frames from multi-camera setups. Experimental results demonstrate that our approach consistently improves rendering quality and semantic segmentation performance, outperforming baseline methods based on random selection and uncertainty-based heuristics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。