SOTA论证挖掘模型其实学的是数据特征,而非论证本质。
Limited Generalizability in Argument Mining: State-Of-The-Art Models Learn Datasets, Not Arguments
- 用17个语料库测试4种Transformer,发现模型依赖词汇线索
- 在陌生数据集上性能大幅下降,泛化能力差
- 联合训练+任务预训练能提升模型鲁棒性
识别论证是自动化话语分析的关键步骤,尤其在政治辩论、网络讨论和科学推理等场景中。尽管理论研究不断深入,实践中的论证挖掘也因大量公开数据集而快速发展。当前基于BERT的模型在各类基准测试中表现优异,被普遍认为具有跨场景适用性。本研究首次大规模重新评估这些SOTA模型在论证识别中的泛化能力。我们在17个英文句级数据集上测试了四种Transformer模型,包括三种标准模型和一种通过对比学习增强的模型。结果表明,模型不同程度地依赖与内容词相关的词汇捷径,意味着其看似进步常源于数据特定线索而非真正理解论证结构。尽管在已知数据集上表现良好,但在未见数据集上性能显著下降。然而,结合任务特定预训练与联合基准训练可有效提升模型的鲁棒性和泛化能力。
原文摘要 · Abstract (English)
Identifying arguments is a necessary prerequisite for various tasks in automated discourse analysis, particularly within contexts such as political debates, online discussions, and scientific reasoning. In addition to theoretical advances in understanding the constitution of arguments, a significant body of research has emerged around practical argument mining, supported by a growing number of publicly available datasets. On these benchmarks, BERT-like transformers have consistently performed best, reinforcing the belief that such models are broadly applicable across diverse contexts of debate. This study offers the first large-scale re-evaluation of such state-of-the-art models, with a specific focus on their ability to generalize in identifying arguments. We evaluate four transformers, three standard and one enhanced with contrastive pre-training for better generalization, on 17 English sentence-level datasets as most relevant to the task. Our findings show that, to varying degrees, these models tend to rely on lexical shortcuts tied to content words, suggesting that apparent progress may often be driven by dataset-specific cues rather than true task alignment. While the models achieve strong results on familiar benchmarks, their performance drops markedly when applied to unseen datasets. Nonetheless, incorporating both task-specific pre-training and joint benchmark training proves effective in enhancing both robustness and generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。