让机器人理解摄影美学并自动找最佳拍摄角度
PhotoAgent: A Robotic Photographer with Spatial and Aesthetic Understanding
- 用大模型推理将抽象美感转化为可计算的几何约束
- 通过3D高斯泼溅构建虚拟世界,快速模拟优化视角
- 适合需要智能摄影决策的机器人或自动化创作场景
用于创意任务(如摄影)的具身智能体必须弥合高层语言指令与几何控制之间的语义鸿沟。我们提出PhotoAgent,该智能体通过整合大模型多模态推理与新型控制范式实现这一目标。PhotoAgent首先利用大模型驱动的思维链推理,将主观美学目标转化为可求解的几何约束,从而由分析求解器计算出高质量初始视角。该初始姿态随后在基于3D高斯泼溅(3DGS)构建的逼真内部世界模型中,通过视觉反馈进行迭代优化。这种‘心理模拟’替代了耗时且昂贵的物理试错过程,实现了快速收敛至更优的视觉结果。评估表明,PhotoAgent在空间推理能力上表现优异,并显著提升了最终图像质量。
原文摘要 · Abstract (English)
Embodied agents for creative tasks like photography must bridge the semantic gap between high-level language commands and geometric control. We introduce PhotoAgent, an agent that achieves this by integrating Large Multimodal Models (LMMs) reasoning with a novel control paradigm. PhotoAgent first translates subjective aesthetic goals into solvable geometric constraints via LMM-driven, chain-of-thought (CoT) reasoning, allowing an analytical solver to compute a high-quality initial viewpoint. This initial pose is then iteratively refined through visual reflection within a photorealistic internal world model built with 3D Gaussian Splatting (3DGS). This ``mental simulation'' replaces costly and slow physical trial-and-error, enabling rapid convergence to aesthetically superior results. Evaluations confirm that PhotoAgent excels in spatial reasoning and achieves superior final image quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。