arXiv:2410.06014cs.ROcs.AI2024-10被引 11

用语言指令控制相机在3D场景中自动寻找最佳拍摄路径。

SplaTraj: Camera Trajectory Generation with Semantic Gaussian Splatting

  • 将相机路径生成转为连续时间优化问题,结合语义理解规划移动轨迹。
  • 通过语言嵌入定位目标区域,使相机自然过渡并拍摄出视觉优美的画面。
  • 适合需要智能拍摄的机器人、虚拟制片或自动化内容生成场景。

近年来,机器人环境建模多聚焦于照片级真实感重建。本文专注于从照片级真实感的高斯点云模型中生成符合用户语言指令的图像序列。我们提出一种新框架 SplaTraj,将环境内图像生成建模为连续时间轨迹优化问题。通过设计代价函数,使沿轨迹运动的相机能平滑穿越环境,并以摄影美感呈现指定的空间信息。该方法利用语言嵌入查询真实感表示,定位与用户输入匹配的区域,再将其投影到随时间变化的相机视角中构建代价。随后通过梯度优化并反向传播渲染过程,求解最优轨迹。最终生成的路径可精准捕捉指定物体的美观视角。我们在多个环境和指令组合上进行评估,验证了生成图像序列的质量。

原文摘要 · Abstract (English)

Many recent developments for robots to represent environments have focused on photorealistic reconstructions. This paper particularly focuses on generating sequences of images from the photorealistic Gaussian Splatting models, that match instructions that are given by user-inputted language. We contribute a novel framework, SplaTraj, which formulates the generation of images within photorealistic environment representations as a continuous-time trajectory optimization problem. Costs are designed so that a camera following the trajectory poses will smoothly traverse through the environment and render the specified spatial information in a photogenic manner. This is achieved by querying a photorealistic representation with language embedding to isolate regions that correspond to the user-specified inputs. These regions are then projected to the camera's view as it moves over time and a cost is constructed. We can then apply gradient-based optimization and differentiate through the rendering to optimize the trajectory for the defined cost. The resulting trajectory moves to photogenically view each of the specified objects. We empirically evaluate our approach on a suite of environments and instructions, and demonstrate the quality of generated image sequences.

相机轨迹语言控制3D生成高斯溅射

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。