用大模型推理实现无需训练的智能物体摆放,支持新对象新场景。
presto: Efficient, Training-free, and Open-world Object Placement via Imaginary Search

- 基于多模态大模型在虚拟空间中迭代优化物体位置与尺寸。
- 在开放世界场景下超越现有方法,人类评估更符合视觉直觉。
- 零样本、免训练,适合快速生成复杂图像布局。
物体摆放对图像构图至关重要,需在多样场景中实现空间与语义上的协调。现有方法依赖人工规则或有限数据集上的监督学习,限制了其泛化能力与可解释性,尤其在包含新对象和新场景的开放世界中表现不佳。本文将开放世界物体摆放重构为由多模态大语言模型(MLLM)引导的启发式搜索任务,提出 extsf{presto} 框架——一种零样本、免训练的方法,在虚构动作空间中迭代优化物体位置与尺度。采用粗到精的搜索策略,确保快速收敛。我们评估了两种决策变体:基于度量的选择与以 MLLM 为裁判的决策。在多个基准测试中, extsf{presto} 实现了领先性能,尤其在未见的开放世界设置中表现突出。人类评估显示,以 MLLM 为裁判的版本比度量驱动方法产生更符合感知一致性的布局,揭示了标准评估指标与人类视觉判断之间的差距。
原文摘要 · Abstract (English)
Object placement is critical in image composition, requiring spatially and semantically coherent positioning of objects within diverse scenes. Existing approaches typically rely on hand-crafted rules or supervised learning on limited datasets, which restricts their generalization and interpretability, especially in open-world scenarios involving novel objects and scenes. In this work, we reformulate open-world object placement as a heuristic search task guided by reasoning from a Multimodal Large Language Model (MLLM). We introduce \textsf{presto}, a zero-shot, training-free framework that operates within an imaginary action space to iteratively refine object position and scale. Our coarse-to-fine search strategy ensures fast convergence, and we evaluate two decision-making variants: Metric-guided Selection and MLLM-as-a-judge. Experiments across multiple benchmarks show that \textsf{presto}~achieves state-of-the-art performance, particularly in previously unseen, open-world settings. Human studies further reveal that the MLLM-as-a-judge variant produces more perceptually coherent placements than metric-driven approaches, highlighting a gap between standard evaluation metrics and human visual judgment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。