arXiv:2605.23771cs.CVcs.AI2026-05

让AI在3D场景中自主拍出符合语义的高质量照片

PhotoFlow: Agentic 3D Virtual Photography Missions

论文配图:PhotoFlow: Agentic 3D Virtual Photography Missions
图 1 · 摘自论文原文
  • 用导演-评审-反思三阶段闭环策略规划相机参数
  • 在6轮内达成最高成功率达78.7%,优于随机搜索等方法
  • 适合对3D生成、智能摄影感兴趣的开发者与研究者

虚拟摄影要求智能体在预设的3D场景中,无初始相机位姿或参考图像,仅根据场景信息和语言指令推断合适构图,选择可执行的相机参数并渲染照片。近期视觉-语言模型的发展使这一空间智能体任务更具可行性,但该任务同时考验复杂3D空间理解与抽象审美判断能力,难以协同评估。我们提出PhotoFlow,一种由导演-评审-反思组成的闭环相机搜索代理:导演构建软性拍摄蓝图并生成多样候选;评审结合规则检查、视觉评价与成对优胜选择;反思模块将失败转化为区域记忆、死区抑制与高探索重定位。我们还构建了VPhotoBench基准,包含47个开源Blender场景和141个语言条件摄影任务,涵盖主体布局、关系构图与氛围风格。在预留测试中,PhotoFlow在六轮渲染预算下,综合外部质量一致性与成功率均优于单次预测、单链反思、锚点池选择和随机搜索。据我们所知,这是首个将任意Blender场景中的语言条件虚拟摄影变为可执行智能体任务的工作,结果表明以大模型为核心的时空智能体已在挑战3D推理与审美选择的设定中生成高质量照片。

原文摘要 · Abstract (English)

Virtual photography asks an agent to enter a prepared 3D scene with no preselected camera pose or reference image, infer a suitable shot from scene information and a language intent, choose executable camera parameters, and render the final photograph. Recent progress in vision-language models makes this kind of spatial agent increasingly plausible, but the task stresses two capabilities that remain hard to evaluate together: complex 3D spatial understanding and abstract aesthetic judgment. We introduce PhotoFlow, a Director-Reviewer-Reflector agent for closed-loop camera search. The Director builds a soft photographic blueprint and proposes diverse candidate cameras; the Reviewer combines rule checks, visual critique, and pairwise incumbent selection; and the Reflector converts failures into region memory, dead-zone suppression, and high-explore relocation. We also introduce VPhotoBench, a benchmark of 47 open-license Blender scenes and 141 language-conditioned photography missions spanning subject placement, relational composition, and atmosphere/style. On held-out experiments, PhotoFlow achieves the strongest external quality-alignment composite and success rate among one-shot prediction, single-chain reflection, anchor-bank selection, and random search under a six-round rendering budget. To our knowledge, this is the first work to make language-conditioned virtual photography in arbitrary Blender scenes an executable agent task, and our results show that an LLM-centered spatial agent can already produce strong photographs in a setting designed to challenge both 3D reasoning and aesthetic choice.

虚拟摄影3D生成智能代理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。