用吉尼指数发现并缓解提示分类中的准确率不公问题
Discovering the Hidden Role of Gini Index In Prompt-based Classification
- 用吉尼指数量化类别准确率差异,识别少数类性能短板
- 在少样本新闻、生物医学和零样本图像任务中显著降低准确率差距
- 无需修改模型即可应用的后处理方法,适合各类提示分类场景
在分类任务中,长尾分布的少数类往往最具实际价值,但其准确率始终偏低,少数高绩效类别主导整体表现。本文深入探究吉尼指数在提示分类中检测与优化(去偏)类别准确率差异的潜在作用。通过真实大模型与视觉模型的基准测试,揭示了吉尼指数不仅是相对准确率主导性的度量工具,更可直接作为优化目标。案例分析表明,无论文本或图像分类、高维或低维输入,均存在从弱到强的相对准确率不平衡现象。基于此,我们提出一种后处理、模型无关的去偏方法。在少样本新闻、生物医学及零样本图像分类任务中,实验结果表明该方法显著降低了相对与绝对准确率不平衡,有效削弱了顶级类别的主导性,同时提升了最弱类别的性能。
原文摘要 · Abstract (English)
In classification tasks, the long-tailed minority classes usually offer the predictions that are most important. Yet these classes consistently exhibit low accuracies, whereas a few high-performing classes dominate the game. We pursue a foundational understanding of the hidden role of Gini Index as a tool for detecting and optimizing (debiasing) disparities in class accuracy, focusing on the case of prompt-based classification. We introduce the intuitions, benchmark Gini scores in real-world LLMs and vision models, and thoroughly discuss the insights of Gini not only as a measure of relative accuracy dominance but also as a direct optimization metric. Through rigorous case analyses, we first show that weak to strong relative accuracy imbalance exists in both prompt-based, text and image classification results and regardless of whether the classification is high-dimensional or low-dimensional. Then, we harness the Gini metric to propose a post-hoc model-agnostic bias mitigation method. Experimental results across few-shot news, biomedical, and zero-shot image classification show that our method significantly reduces both relative and absolute accuracy imbalances, minimizing top class relative dominance while elevating weakest classes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。