用自然语言指令实现多摄像头图像智能拼接,解决盲区与畸变问题。
ChatStitch: Visualizing Through Structures via Surround-View Unsupervised Deep Image Stitching with Collaborative LLM-Agents
- 基于大模型的多智能体闭环交互,支持自然语言控制拼接过程。
- 在3/4/5图拼接任务中PSNR提升9%/17%/21%,SSIM提升8%/18%/26%。
- 适合自动驾驶、智能监控等需多视角融合的场景。
环绕视图感知因其能通过周围摄像头的信息交换增强自动驾驶车辆的感知能力而受到广泛关注。然而,现有系统受限于与人类的单向交互模式以及重叠区域畸变随非重叠区域指数级传播的问题。为此,本文提出ChatStitch,一个支持通过自然语言指令与外部数字资产协同的环绕视图人机共感知系统,可揭示被遮挡的盲区信息。为打破单向交互瓶颈,ChatStitch构建了基于大语言模型的认知闭环多智能体框架。为抑制重叠边界处畸变传播,提出SV-UDIS方法,在非全局重叠条件下实现无监督深度图像拼接。我们在UDIS-D、MCOV-SLAM公开数据集及自建真实世界数据集上进行了大量实验。结果表明,该方法在UDIS-D数据集上3、4、5图拼接任务中均达到当前最优性能,PSNR分别提升9%、17%、21%,SSIM分别提升8%、18%、26%。
原文摘要 · Abstract (English)
Surround-view perception has garnered significant attention for its ability to enhance the perception capabilities of autonomous driving vehicles through the exchange of information with surrounding cameras. However, existing surround-view perception systems are limited by inefficiencies in unidirectional interaction pattern with human and distortions in overlapping regions exponentially propagating into non-overlapping areas. To address these challenges, this paper introduces ChatStitch, a surround-view human-machine co-perception system capable of unveiling obscured blind spot information through natural language commands integrated with external digital assets. To dismantle the unidirectional interaction bottleneck, ChatStitch implements a cognitively grounded closed-loop interaction multi-agent framework based on Large Language Models. To suppress distortion propagation across overlapping boundaries, ChatStitch proposes SV-UDIS, a surround-view unsupervised deep image stitching method under the non-global-overlapping condition. We conducted extensive experiments on the UDIS-D, MCOV-SLAM open datasets, and our real-world dataset. Specifically, our SV-UDIS method achieves state-of-the-art performance on the UDIS-D dataset for 3, 4, and 5 image stitching tasks, with PSNR improvements of 9\%, 17\%, and 21\%, and SSIM improvements of 8\%, 18\%, and 26\%, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。