构建首个大语言模型数据归属评估基准,助力模型可解释性研究。
DATE-LM: Benchmarking Data Attribution Evaluation for Large Language Models
- 设计三类真实任务评估数据贡献度,覆盖训练数据筛选、毒性过滤等场景。
- 实测发现无方法在所有任务中领先,且性能受任务设计影响显著。
- 提供公开排行榜,推动社区共建数据归属评估体系。
数据归属方法用于量化训练数据对模型输出的影响,日益重要于数据集优化、模型可解释性与数据估值等领域。然而,现有研究缺乏系统性的大语言模型中心评估。为此,我们提出 DATE-LM(大语言模型中的数据归属评估),一个统一基准,通过真实世界大语言模型应用评估数据归属方法。该基准涵盖三项关键任务:训练数据选择、毒性/偏见过滤与事实归属。其设计简洁易用,支持跨任务、多模型架构的大规模评估。我们利用 DATE-LM 对现有方法进行大规模测评,结果表明:单一方法无法在所有任务中胜出,数据归属方法常不如简单基线,且性能高度依赖任务设计。最后,我们发布公开排行榜,促进方法比较与社区协作,旨在使 DATE-LM 成为未来大语言模型数据归属研究的基石。
原文摘要 · Abstract (English)
Data attribution methods quantify the influence of training data on model outputs and are becoming increasingly relevant for a wide range of LLM research and applications, including dataset curation, model interpretability, data valuation. However, there remain critical gaps in systematic LLM-centric evaluation of data attribution methods. To this end, we introduce DATE-LM (Data Attribution Evaluation in Language Models), a unified benchmark for evaluating data attribution methods through real-world LLM applications. DATE-LM measures attribution quality through three key tasks -- training data selection, toxicity/bias filtering, and factual attribution. Our benchmark is designed for ease of use, enabling researchers to configure and run large-scale evaluations across diverse tasks and LLM architectures. Furthermore, we use DATE-LM to conduct a large-scale evaluation of existing data attribution methods. Our findings show that no single method dominates across all tasks, data attribution methods have trade-offs with simpler baselines, and method performance is sensitive to task-specific evaluation design. Finally, we release a public leaderboard for quick comparison of methods and to facilitate community engagement, with the motivation that DATE-LM can serve as a foundation for future data attribution research in LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。