arXiv:2509.17757cs.CVcs.MA2025-09被引 6

多智能体协作生成被遮挡物体的完整形态,效果更准更稳。

Multi-Agent Amodal Completion: Direct Synthesis with Fine-Grained Semantic Guidance

  • 用多个智能体协同分析遮挡关系并生成精准填充区域
  • 直接输出带透明度的图层,无需额外分割步骤
  • 适合需要高精度图像修复与增强的视觉应用

可见物体的不完整部分生成(amodal completion)对图像编辑和增强现实等应用至关重要。现有方法存在数据依赖强、泛化能力差或渐进式流程中误差累积的问题。本文提出基于前置协作推理的多智能体协同框架:多个智能体协同分析遮挡关系,确定边界扩展范围,生成精确掩码用于修复;同时,一个智能体生成细粒度文本描述,提供细粒度语义引导,确保合成准确,避免重新生成遮挡物或其他无关元素,尤其在大范围修复时表现更优。此外,该方法直接输出由扩散变换器生成的可见掩码和注意力图引导的分层RGBA结果,无需额外分割步骤。大量实验表明,本方法在视觉质量上达到当前最优水平。

原文摘要 · Abstract (English)

Amodal completion, generating invisible parts of occluded objects, is vital for applications like image editing and AR. Prior methods face challenges with data needs, generalization, or error accumulation in progressive pipelines. We propose a Collaborative Multi-Agent Reasoning Framework based on upfront collaborative reasoning to overcome these issues. Our framework uses multiple agents to collaboratively analyze occlusion relationships and determine necessary boundary expansion, yielding a precise mask for inpainting. Concurrently, an agent generates fine-grained textual descriptions, enabling Fine-Grained Semantic Guidance. This ensures accurate object synthesis and prevents the regeneration of occluders or other unwanted elements, especially within large inpainting areas. Furthermore, our method directly produces layered RGBA outputs guided by visible masks and attention maps from a Diffusion Transformer, eliminating extra segmentation. Extensive evaluations demonstrate our framework achieves state-of-the-art visual quality.

图像修复多智能体扩散模型语义引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。