少样本关系分类中,关系多样性比数据量更重要。
Diversity Over Quantity: A Lesson From Few Shot Relation Classification
- 用更多样化的关系类型训练模型,提升泛化能力。
- 在固定数据量下,关系多样性提升使准确率显著提高。
- 适合关注小样本学习与数据高效设计的研究者。
在少样本关系分类(FSRC)中,模型需仅凭少量标注样本就泛化到新关系。尽管近期自然语言处理进展多聚焦于扩大数据规模,我们提出:关系类型的多样性对FSRC性能更为关键。本文通过系统实验表明,在训练数据总量不变的情况下,增加关系类型的多样性能持续提升模型在多种少样本场景下的表现,包括高负例设置。为此,我们构建了REBEL-FS,一个关系类型数量比现有数据集多一个数量级的新基准。研究结果挑战了“数据越多越好”的普遍假设,表明聚焦多样性的数据筛选可大幅减少对大规模数据集的依赖。
原文摘要 · Abstract (English)
In few-shot relation classification (FSRC), models must generalize to novel relations with only a few labeled examples. While much of the recent progress in NLP has focused on scaling data size, we argue that diversity in relation types is more crucial for FSRC performance. In this work, we demonstrate that training on a diverse set of relations significantly enhances a model's ability to generalize to unseen relations, even when the overall dataset size remains fixed. We introduce REBEL-FS, a new FSRC benchmark that incorporates an order of magnitude more relation types than existing datasets. Through systematic experiments, we show that increasing the diversity of relation types in the training data leads to consistent gains in performance across various few-shot learning scenarios, including high-negative settings. Our findings challenge the common assumption that more data alone leads to better performance and suggest that targeted data curation focused on diversity can substantially reduce the need for large-scale datasets in FSRC.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。