用LLM答题差值判断文本集是否有信息价值,省去训练成本。
Is This Collection Worth My LLM's Time? Automatically Measuring Information Potential in Text Corpora
- 生成选择题测试LLM在有无原文时的答题表现差异。
- 五类数据集验证,能有效识别含新知识的文本集。
- 适合想选优质数据但缺资源的团队使用。
随着大语言模型(LLMs)能力趋于趋同,提升性能的关键在于发现并整合有价值的新信息源。然而,评估哪些文本集合值得投入大量资源进行数字化、预处理及集成到LLM系统中仍具挑战性。本文提出一种无需模型训练或微调的自动化评估方法:从文本生成多选题(MCQs),测量LLM在有无原始文本支持下的表现差异,该差距作为信息潜力的代理指标。我们在五个精心选取的数据集上验证该方法:EPFL博士论文、一份威尼斯历史档案私藏、两组相关主题的维基百科文章,以及一个合成基准数据集。结果表明,该方法能有效识别包含新颖信息的文本集合,为数据获取与集成优先级提供实用工具。
原文摘要 · Abstract (English)
As large language models (LLMs) converge towards similar capabilities, the key to advancing their performance lies in identifying and incorporating valuable new information sources. However, evaluating which text collections are worth the substantial investment required for digitization, preprocessing, and integration into LLM systems remains a significant challenge. We present a novel approach to this challenge: an automated pipeline that evaluates the potential information gain from text collections without requiring model training or fine-tuning. Our method generates multiple choice questions (MCQs) from texts and measures an LLM's performance both with and without access to the source material. The performance gap between these conditions serves as a proxy for the collection's information potential. We validate our approach using five strategically selected datasets: EPFL PhD manuscripts, a private collection of Venetian historical records, two sets of Wikipedia articles on related topics, and a synthetic baseline dataset. Our results demonstrate that this method effectively identifies collections containing valuable novel information, providing a practical tool for prioritizing data acquisition and integration efforts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。