arXiv:2608.30297cs.CL2026-08

无需标注子组,自动增强不平衡数据以提升模型鲁棒性。

AIA$^{2}$: Attribute-Agnostic Imbalance Augmentation for Subgroup Robustness

论文配图:AIA$^{2}$: Attribute-Agnostic Imbalance Augmentation for Subgroup Robustness
图 1 · 摘自论文原文
  • 通过潜在语义分布自动发现子组不平衡模式
  • 在低表现子组上实现显著性能提升,优于现有基线
  • 适合关注公平性与鲁棒性的自然语言处理研究者

描述数据内容和上下文的属性可能引发超出标签不平衡的多样化不平衡模式。现有研究主要关注标签不平衡,忽略了话题、人口统计等属性带来的有意义子组结构及对少数子组的模型退化问题。本文提出属性无关的不平衡增强框架AIA²,无需显式子组标注即可提升模型在不同子组不平衡下的鲁棒性。AIA²通过潜在语义分布自动识别不平衡,提取学习难度高且不平衡严重的数据片段,并利用大语言模型进行子组感知的增强。在5个涵盖社会议题与多样化主题的主流语料库上评估,结果表明其在表现最差的子组上性能显著提升,且持续优于竞争性基线。消融实验验证各组件互补贡献,额外分析显示AIA²是提升最差子组鲁棒性的实用且一致的方法。代码已开源。

原文摘要 · Abstract (English)

Attributes describing data content and context can induce diverse imbalance patterns that go beyond label imbalance alone. However, existing studies primarily address label imbalance while overlooking data attributes, such as topics and demographics, which can induce meaningful subgroup structure while causing model degradation on underrepresented subgroups. We propose Attribute-Agnostic Imbalance Augmentation (AIA$^{2}$), a framework for improving model robustness under varying subgroup imbalances without explicit subgroup annotations. AIA$^{2}$ automatically discovers varying imbalances via latent semantic distributions, obtains slices with both learning difficulty and subgroup imbalance deficits, and deploys a large language model (LLM) for subgroup-aware imbalance augmentation. We have evaluated AIA$^{2}$ on 5 popular corpora with rich domains and their attribute values, covering social issues and diverse topics. Results show improved performance on the lowest-performing subgroups and consistent gains over competitive baselines. Ablation studies confirm complementary contributions from each component, and additional analyses show that AIA$^{2}$ provides a practical and consistent way to improve worst-group robustness under data subgroup imbalance. Code is available at https://github.com/trust-nlp/AIA2-Subgroup-Robustness.

子组鲁棒性数据不平衡公平性LLM增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。