arXiv:2506.21839cs.CVcs.CL2025-06ICCV

用多智能体协作生成逻辑严密的密室逃脱谜题图像

GenEscape: Hierarchical Multi-Agent Generation of Escape Room Puzzles

  • 分阶段构建:功能设计→符号场景图→布局合成→局部编辑
  • 协作智能体使谜题可解率提升,避免捷径,清晰表达物品用途
  • 适合游戏设计、创意生成与具身智能研究者参考

我们挑战文本到图像模型,生成视觉吸引人、逻辑严谨且富有智力挑战性的密室逃脱谜题图像。基础图像模型在空间关系和物体功能推理上表现不足,为此提出分层多智能体框架,将任务分解为功能设计、符号场景图推理、布局合成和局部图像编辑四个结构化阶段。专用智能体通过迭代反馈协作,确保场景视觉连贯且功能可解。实验表明,智能体协作提升了输出质量,在可解性、避免捷径和功能清晰度方面均有改善,同时保持良好视觉效果。

原文摘要 · Abstract (English)

We challenge text-to-image models with generating escape room puzzle images that are visually appealing, logically solid, and intellectually stimulating. While base image models struggle with spatial relationships and affordance reasoning, we propose a hierarchical multi-agent framework that decomposes this task into structured stages: functional design, symbolic scene graph reasoning, layout synthesis, and local image editing. Specialized agents collaborate through iterative feedback to ensure the scene is visually coherent and functionally solvable. Experiments show that agent collaboration improves output quality in terms of solvability, shortcut avoidance, and affordance clarity, while maintaining visual quality.

密室逃脱多智能体图像生成逻辑推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。