arXiv:2512.10178cs.LGcs.CL2025-12被引 1

通过聚类控制生成方向,补全数据分布盲区。

CIEGAD: Cluster-Conditioned Interpolative and Extrapolative Framework for Geometry-Aware and Domain-Aligned Data Augmentation

  • 用聚类条件控制生成,分层分配频率与几何特征。
  • 在长尾和多分类任务中提升F1与召回率。
  • 适合数据稀缺或类别不平衡场景使用。

实际深度学习部署中,数据稀缺与标签分布不均常导致真实数据分布中存在语义未覆盖区域,影响模型训练,并引发类别边界附近的误分类及边缘区域的不稳定行为。尽管大语言模型(LLMs)在数据增强方面展现潜力,但兼具生成方向控制、领域对齐与质量控制的集成框架尚未成熟。为此,我们提出面向几何感知与领域对齐的数据增强框架CIEGAD,系统性补全分布内与分布外的语义盲区。CIEGAD通过聚类条件构建领域画像,采用融合类别频率与几何指标的分层频率-几何分配策略进行生成分配,并通过插值与外推合成共存实现生成方向精细控制。同时,结合几何约束过滤与基于LLM的判别机制进行质量控制。多个分类任务实验表明,CIEGAD有效扩展了真实数据分布的外围,保持生成数据与真实数据的高度对齐性与语义多样性。尤其在长尾与多分类任务中,持续提升F1与召回率,验证了分布一致性、多样性与质量的三重协调。结果表明,CIEGAD是一个面向实践的数据增强框架,可在补全低频区域的同时保持与真实数据的对齐。

原文摘要 · Abstract (English)

In practical deep learning deployment, the scarcity of data and the imbalance of label distributions often lead to semantically uncovered regions within the real-world data distribution, hindering model training and causing misclassification near class boundaries as well as unstable behaviors in peripheral areas. Although recent large language models (LLMs) show promise for data augmentation, an integrated framework that simultaneously achieves directional control of generation, domain alignment, and quality control has not yet been fully established. To address these challenges, we propose a Cluster-conditioned Interpolative and Extrapolative framework for Geometry-Aware and Domain-aligned data augmentation (CIEGAD), which systematically complements both in-distribution and out-of-distribution semantically uncovered regions. CIEGAD constructs domain profiles through cluster conditioning, allocates generation with a hierarchical frequency-geometric allocation integrating class frequency and geometric indicators, and finely controls generation directions via the coexistence of interpolative and extrapolative synthesis. It further performs quality control through geometry-constrained filtering combined with an LLM-as-a-Judge mechanism. Experiments on multiple classification tasks demonstrate that CIEGAD effectively extends the periphery of real-world data distributions while maintaining high alignment between generated and real-world data as well as semantic diversity. In particular, for long-tailed and multi-class classification tasks, CIEGAD consistently improves F1 and recall, validating the triple harmony of distributional consistency, diversity, and quality. These results indicate that CIEGAD serves as a practically oriented data augmentation framework that complements underrepresented regions while preserving alignment with real-world data.

数据增强长尾分布生成模型几何感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。