arXiv:2503.22172cs.CV2025-03中稿 · CVPR被引 1

用概念感知的LoRA微调生成更真实的城市场景分割数据集

CA-LoRA: Concept-Aware LoRA for Domain-Aligned Segmentation Dataset Generation

  • 仅更新与风格、视角等关键概念相关的权重,保留预训练知识
  • 在少样本和全监督设置下均优于基线与最新方法
  • 特别适合生成恶劣天气和光照变化下的多样化数据

本文针对语义分割中的数据稀缺问题,提出通过文本到图像(T2I)生成模型自动生成数据集,以降低图像采集与标注成本。数据集生成面临两大挑战:一是使生成样本与目标领域对齐,二是生成超越训练数据的信息丰富样本。微调T2I模型有助于实现领域对齐,但常导致过拟合与记忆训练数据,限制多样性。为此,我们提出概念感知的LoRA(CA-LoRA),一种新型微调方法,仅选择性地更新与必要概念(如风格或视角)相关的权重,同时保留T2I模型的预训练知识,从而生成信息丰富且对齐良好的样本。我们在城市场景分割数据集生成中验证了其有效性,在域内(少样本与全监督)及域泛化任务中均优于基线与最先进方法,尤其在恶劣天气和光照变化条件下表现更优,充分体现了其优势。

原文摘要 · Abstract (English)

This paper addresses the challenge of data scarcity in semantic segmentation by generating datasets through text-to-image (T2I) generation models, reducing image acquisition and labeling costs. Segmentation dataset generation faces two key challenges: 1) aligning generated samples with the target domain and 2) producing informative samples beyond the training data. Fine-tuning T2I models can help generate samples aligned with the target domain. However, it often overfits and memorizes training data, limiting their ability to generate diverse and well-aligned samples. To overcome these issues, we propose Concept-Aware LoRA (CA-LoRA), a novel fine-tuning approach that selectively identifies and updates only the weights associated with necessary concepts (e.g., style or viewpoint) for domain alignment while preserving the pretrained knowledge of the T2I model to produce informative samples. We demonstrate its effectiveness in generating datasets for urban-scene segmentation, outperforming baseline and state-of-the-art methods in in-domain (few-shot and fully-supervised) settings, as well as in domain generalization tasks, especially under challenging conditions such as adverse weather and varying illumination, further highlighting its superiority.

图像生成分割数据集LoRA域对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。