通过模型知识对比,量化不同数据子集间的相似性。
RepMatch: Quantifying Cross-Instance Similarities in Representation Space
- 用模型编码的知识比较不同数据子集的相似性。
- 能选出比随机抽样表现更好的代表性数据子集。
- 适合研究数据结构、评估数据质量的NLP研究人员。
数据集分析技术的进步使得对训练数据实例进行更复杂的分析和表征成为可能,通常基于诸如“难度”等属性对数据进行分类。本文提出RepMatch,一种通过相似性视角刻画数据的新方法。该方法通过比较在不同数据子集上训练的模型所编码的知识,量化这些子集之间的相似性,克服了现有分析方法仅关注单个实例且局限于同一数据集内分析的局限。该框架支持任意数据子集间的相似性比较,可实现跨数据集与实例级的分析。我们在多个NLP任务、数据集和模型上验证了RepMatch的有效性。大量实验表明,RepMatch能够有效比较不同数据集,识别出更具代表性的数据子集(其性能优于同规模随机选取的子集),并揭示部分挑战性数据集构建背后的启发式规律。
原文摘要 · Abstract (English)
Advances in dataset analysis techniques have enabled more sophisticated approaches to analyzing and characterizing training data instances, often categorizing data based on attributes such as ``difficulty''. In this work, we introduce RepMatch, a novel method that characterizes data through the lens of similarity. RepMatch quantifies the similarity between subsets of training instances by comparing the knowledge encoded in models trained on them, overcoming the limitations of existing analysis methods that focus solely on individual instances and are restricted to within-dataset analysis. Our framework allows for a broader evaluation, enabling similarity comparisons across arbitrary subsets of instances, supporting both dataset-to-dataset and instance-to-dataset analyses. We validate the effectiveness of RepMatch across multiple NLP tasks, datasets, and models. Through extensive experimentation, we demonstrate that RepMatch can effectively compare datasets, identify more representative subsets of a dataset (that lead to better performance than randomly selected subsets of equivalent size), and uncover heuristics underlying the construction of some challenge datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。