用稀疏视角生成3D美学场,高效推荐美观拍摄位置。
Aesthetic Camera Viewpoint Suggestion with 3D Aesthetic Field
- 构建3D美学场,从稀疏图像推断全局美学分布。
- 相比强化学习,搜索效率提升显著,推荐视角更优。
- 适合摄影、虚拟拍摄等需3D美学判断的场景。
场景的美学质量高度依赖于相机视角。现有方法或仅支持单视角调整,无法理解场景几何;或依赖密集采集与预建3D环境,需代价高昂的强化学习搜索。本文提出3D美学场概念,基于稀疏视角实现几何感知的美学推理,无需密集重建或强化学习。我们采用前馈式3D高斯泼溅网络,将预训练2D美学模型的高层美学知识蒸馏至3D空间,仅凭稀疏输入视图即可预测新视角的美学评分。在此基础上,设计两阶段搜索流程:先粗采样后梯度优化,高效识别美学上佳视角。大量实验表明,本方法在构图与取景上优于现有方法,为3D感知美学建模开辟新路径。
原文摘要 · Abstract (English)
The aesthetic quality of a scene depends strongly on camera viewpoint. Existing approaches for aesthetic viewpoint suggestion are either single-view adjustments, predicting limited camera adjustments from a single image without understanding scene geometry, or 3D exploration approaches, which rely on dense captures or prebuilt 3D environments coupled with costly reinforcement learning (RL) searches. In this work, we introduce the notion of 3D aesthetic field that enables geometry-grounded aesthetic reasoning in 3D with sparse captures, allowing efficient viewpoint suggestions in contrast to costly RL searches. We opt to learn this 3D aesthetic field using a feedforward 3D Gaussian Splatting network that distills high-level aesthetic knowledge from a pretrained 2D aesthetic model into 3D space, enabling aesthetic prediction for novel viewpoints from only sparse input views. Building on this field, we propose a two-stage search pipeline that combines coarse viewpoint sampling with gradient-based refinement, efficiently identifying aesthetically appealing viewpoints without dense captures or RL exploration. Extensive experiments show that our method consistently suggests viewpoints with superior framing and composition compared to existing approaches, establishing a new direction toward 3D-aware aesthetic modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。