用自然语言解释文本分类器的系统性错误并提升性能
DISCERN: Decoding Systematic Errors in Natural Language for Text Classifiers
- 双大模型交互生成错误的精准语言描述
- 语言解释使分类器性能提升,优于仅用反例
- 人类理解错误更高效,效率提升超25%
尽管当前机器学习系统预测准确率高,但常因标注缺陷或某些类别数据不足而产生系统性偏差。现有方法尝试通过关键词自动识别和解释此类偏差。本文提出DISCERN框架,利用两个大语言模型的交互循环,生成文本分类器系统性错误的精确自然语言描述。随后,基于这些描述,通过主动学习或合成数据增强训练集以改进分类器。在三个文本分类数据集上,实验表明语言解释带来的性能提升具有一致性,且超越仅使用偏差样例的效果。人类评估显示,相比聚类样例,通过语言解释理解系统性偏差的效率和有效性均提升超过25%。
原文摘要 · Abstract (English)
Despite their high predictive accuracies, current machine learning systems often exhibit systematic biases stemming from annotation artifacts or insufficient support for certain classes in the dataset. Recent work proposes automatic methods for identifying and explaining systematic biases using keywords. We introduce DISCERN, a framework for interpreting systematic biases in text classifiers using language explanations. DISCERN iteratively generates precise natural language descriptions of systematic errors by employing an interactive loop between two large language models. Finally, we use the descriptions to improve classifiers by augmenting classifier training sets with synthetically generated instances or annotated examples via active learning. On three text-classification datasets, we demonstrate that language explanations from our framework induce consistent performance improvements that go beyond what is achievable with exemplars of systematic bias. Finally, in human evaluations, we show that users can interpret systematic biases more effectively (by over 25% relative) and efficiently when described through language explanations as opposed to cluster exemplars.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。