分析语言模型在文本分类中依赖复杂捷径的问题
Navigating the Shortcut Maze: A Comprehensive Analysis of Shortcut Learning in Text Classification by Language Models
- 按出现、风格、概念三类构建文本分类捷径基准
- 发现主流模型对复杂捷径仍高度敏感,泛化能力受限
- 适合关注模型可靠性与鲁棒性的研究者参考
尽管语言模型(LMs)取得显著进展,但其常依赖虚假相关性,影响准确性和泛化能力。本研究关注被忽视的更隐蔽、更复杂的捷径对模型可靠性的影响。我们提出一个综合性基准,将捷径分为出现、风格和概念三类,系统探究这些复杂捷径如何影响语言模型的性能。通过在传统语言模型、大语言模型及前沿鲁棒模型上的广泛实验,研究揭示了模型在应对复杂捷径时的脆弱性与韧性。相关基准与代码已公开于:https://github.com/yuqing-zhou/shortcut-learning-in-text-classification。
原文摘要 · Abstract (English)
Language models (LMs), despite their advances, often depend on spurious correlations, undermining their accuracy and generalizability. This study addresses the overlooked impact of subtler, more complex shortcuts that compromise model reliability beyond oversimplified shortcuts. We introduce a comprehensive benchmark that categorizes shortcuts into occurrence, style, and concept, aiming to explore the nuanced ways in which these shortcuts influence the performance of LMs. Through extensive experiments across traditional LMs, large language models, and state-of-the-art robust models, our research systematically investigates models' resilience and susceptibilities to sophisticated shortcuts. Our benchmark and code can be found at: https://github.com/yuqing-zhou/shortcut-learning-in-text-classification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。