arXiv:2606.18555cs.CV2026-06

用AI生成室内图扩充数据,提升识别准确率

Rethinking Text-to-Image as Semantic-Aware Data Augmentation for Indoor Scene Recognition

论文配图:Rethinking Text-to-Image as Semantic-Aware Data Augmentation for Indoor Scene Recognition
图 1 · 摘自论文原文
  • 用Stable Diffusion生成逼真室内场景图作为数据增强
  • 在MIT Indoor数据集上显著提升模型性能,有限数据下效果更优
  • 提出新方法可100%识别生成图像,适合小模型部署

室内图像识别因光照变化、遮挡和物体布局复杂而面临挑战。为缓解真实训练图像不足的问题,我们提出一种基于Stable Diffusion(SD)生成合成图像的新方法,作为强大的数据增强手段。该方法构建了系统化框架,生成多样化且逼真的室内场景,有效扩充训练数据。在MIT Indoor Scene数据集上的实验表明,该方法在真实数据有限时能显著提升深度模型的训练效果。为进一步防范合成图像滥用,我们引入基于DIffusion Reconstruction Error(DIRE)的检测机制。DIRE表征使仅使用轻量级深度模型即可训练出鲁棒分类器,实验显示使用MobilenetV3可实现100%准确率识别生成图像。

原文摘要 · Abstract (English)

In the realm of computer vision, indoor image recognition presents challenges due to the intricate interplay of lighting conditions, occlusions, and diverse object arrangements within confined spaces. To address the lacks of training indoor images, we introduce a novel approach leveraging Stable Diffusion (SD) for the generation of synthetic images, which serve as a powerful data augmentation tool. The utilization of SD offers a principled framework for synthesizing diverse and realistic indoor scenes, thereby enriching the training data pool for robust indoor image recognition models. Experimental findings on the MIT Indoor Scene dataset reveal the potential of our proposed approach in enhancing the training of deep models when authentic data is limited. Furthermore, to prevent the misuse of SD synthetic images, we introduce a counter measure based on DIffusion Reconstruction Error (DIRE). The powerful DIRE presentation enables training robust classifiers only using lightweight deep models. Experiments show that our approach can perfectly recognize SD generated images with the accuracy of 100% using MobilenetV3.

图像生成数据增强室内识别扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。