arXiv:2602.02841cs.LGcs.AI2026-02

用语义感知生成模型,在数据少时也能高效合成高质量样本。

Semantics-Aware Generative Latent Data Augmentation for Learning in Low-Resource Domains

  • 在大模型隐空间中,用条件扩散模型生成新样本
  • 零样本语音情感识别提升6.13%召回率,长尾图像分类达74.7%准确率
  • 适合小样本、类别不平衡场景下的模型训练

尽管深度学习在数据丰富环境下表现优异,但在实际中常见的数据稀缺场景下性能下降。虽然在大规模数据上训练的基础模型(FMs)能提取通用特征并具备强泛化能力,但在下游微调时仍受限于标注数据不足。为此,我们提出GeLDA——一种语义感知的生成式隐空间数据增强框架,利用条件扩散模型在FM诱导的隐空间中合成样本。由于该空间维度低且集中了任务相关的信息,相比输入空间,可实现高效高质量的数据生成。GeLDA通过辅助特征向量对生成过程施加语义约束,捕捉类别或子域间的语义关系,从而支持低资源领域的数据增强。我们在两个大规模识别任务中验证了GeLDA:(a) 零样本语言特定语音情感识别中,使Whisper-large基线的未加权平均召回率提升6.13%;(b) 在长尾图像分类任务中,在ImageNet-LT上达到74.7%的尾部类别准确率,创下新纪录。

原文摘要 · Abstract (English)

Despite strong performance in data-rich regimes, deep learning often underperforms in the data-scarce settings common in practice. While foundation models (FMs) trained on massive datasets demonstrate strong generalization by extracting general-purpose features, they can still suffer from scarce labeled data during downstream fine-tuning. To address this, we propose GeLDA, a semantics-aware generative latent data augmentation framework that leverages conditional diffusion models to synthesize samples in an FM-induced latent space. Because this space is low-dimensional and concentrates task-relevant information compared to the input space, GeLDA enables efficient, high-quality data generation. GeLDA conditions generation on auxiliary feature vectors that capture semantic relationships among classes or subdomains, facilitating data augmentation in low-resource domains. We validate GeLDA in two large-scale recognition tasks: (a) in zero-shot language-specific speech emotion recognition, GeLDA improves the Whisper-large baseline's unweighted average recall by 6.13%; and (b) in long-tailed image classification, it achieves 74.7% tail-class accuracy on ImageNet-LT, setting a new state-of-the-art result.

数据增强低资源学习扩散模型语义感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。