arXiv:2505.12533cs.CL2025-05中稿 · as a main conferen…被引 2

语言模型在人物关系抽取中泛化能力差,依赖数据噪声而非真实模式。

Relation Extraction or Pattern Matching? Unravelling the Generalisation Limits of Language Models for Biographical RE

  • 通过跨数据集实验发现模型易受特定数据集特征干扰。
  • 高质量数据下微调效果最好,低质量数据下提示学习更优。
  • 现有评测基准结构缺陷限制了模型迁移能力,适合评估鲁棒性研究者参考。

分析关系抽取(RE)模型的泛化能力对于判断其是否学习到稳健的语义模式至关重要。我们的跨数据集实验发现,即使在相似领域内,现有模型也难以适应新数据。值得注意的是,模型在单个数据集上的高性能并不意味着更好的迁移能力,反而常反映对数据集特有噪声的过拟合。结果表明,数据质量比词汇相似度更关键;当数据质量高时,微调可实现最佳跨数据集性能;而面对噪声数据时,少样本上下文学习(ICL)表现更优。然而,即便如此,零样本基线仍偶尔优于所有跨数据集结果。此外,当前评测基准存在的结构性问题——如每样本仅含单一关系、负类定义不统一——进一步阻碍了模型的迁移表现。

原文摘要 · Abstract (English)

Analysing the generalisation capabilities of relation extraction (RE) models is crucial for assessing whether they learn robust relational patterns or rely on spurious correlations. Our cross-dataset experiments find that RE models struggle with unseen data, even within similar domains. Notably, higher intra-dataset performance does not indicate better transferability, instead often signaling overfitting to dataset-specific artefacts. Our results also show that data quality, rather than lexical similarity, is key to robust transfer, and the choice of optimal adaptation strategy depends on the quality of data available: while fine-tuning yields the best cross-dataset performance with high-quality data, few-shot in-context learning (ICL) is more effective with noisier data. However, even in these cases, zero-shot baselines occasionally outperform all cross-dataset results. Structural issues in RE benchmarks, such as single-relation per sample constraints and non-standardised negative class definitions, further hinder model transferability.

关系抽取模型泛化数据质量少样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。