arXiv:2607.03046cs.CV2026-07

用文本生成投影图像,提升摄像头与投影仪对齐精度

Text-to-Image Generation for Projector-Camera System Registration

论文配图:Text-to-Image Generation for Projector-Camera System Registration
图 1 · 摘自论文原文
  • 基于文本生成自然且富含空间特征的投影图像
  • 在多种配置下显著提升投影-摄像机配准准确率
  • 兼顾视觉自然性与定位精度,适合实际应用

在投影-相机系统中建立投影与相机图像间的对应关系(即系统标定)是实现高分辨率像素匹配的关键。传统结构光方法(如条纹或斑点)虽精度高,但效率低且对人眼无意义。现有自然图像方法因特征分布不足,难以达到同等精度。此外,环境因素如表面纹理、光照变化等也影响性能。为此,我们提出一种基于深度神经网络的方法:从文本提示生成单一自然图像,既具真实感又具备丰富空间特征,以提升标定精度。模型在模拟了几何与光照畸变的合成数据集上训练,能有效预测投影与相机图像的对应关系,在多种系统配置下显著提高配准精度。该方法同时保证投影内容视觉自然,减少干扰,用户研究证实其在感知自然度和可用性上优于现有方法,具备实际应用价值。

原文摘要 · Abstract (English)

Establishing correspondence between projector and camera images in a procam (projector + camera) system is essential for achieving high-resolution pixel matching, referred to as procam registration. The highest accuracy is typically obtained using structured light patterns (e.g., stripes or blobs). However, these methods are often inefficient and lack meaningful information for human viewers. Although some have explored the use of natural images, these often fail to provide a sufficient distribution of features to achieve comparable accuracy. Additionally, existing methods struggle to cope with environmental factors such as surface textures and variations in brightness due to ambient light or changes in camera exposure. To address these limitations, we propose a method based on deep neural networks. Our approach aims to generate a single natural image from text-based prompts that not only appears realistic but also possesses rich spatial features to enhance registration accuracy in procam applications. We have developed a deep neural network trained on a synthesized dataset that simulates potential geometric and photometric distortions encountered in a procam system illuminating a relatively smooth object (see Figure 1). Our trained network predicts the correspondence between projector and camera images, significantly improving registration accuracy across various procam configurations. By jointly considering the naturalness and feature richness of the projector images, our method minimizes visual disruptions in projected content without sacrificing precision. A user study confirms that our technique enhances perceived naturalness and usability compared to existing methods, validating its practical utility in real-world applications.

图像生成系统标定深度学习投影校准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。