私有评估隐含利益冲突,可能扭曲大模型真实性能判断
Peeking Behind Closed Doors: Risks of LLM Evaluation by Private Data Curators
- 私有评估由企业关联的专家团队进行,易受商业关系影响
- 专家主观偏好使自家数据训练的模型获得不公平优势
- 适合关注模型评测公平性与行业透明度的研究者
大型语言模型(LLM)的快速发展加剧了科技巨头与初创公司之间的竞争。在此背景下,模型评估对产品决策和投资选择至关重要。尽管开放评估集如MMLU曾推动技术进步,但数据污染和偏见问题持续引发对其可靠性的质疑。因此,私人数据提供方开始使用自建高质量测试题和专家标注员开展隐蔽评估。本文指出,尽管私有评估在缓解数据污染方面有一定优势,却引入了意外的财务与评估风险。主要问题包括:数据提供方与其客户(主流大模型公司)之间的业务关系可能导致利益冲突;此外,专家标注员的主观偏好会带来评估偏差,使使用其数据训练的模型获得优势。本文为研究私有评估风险奠定基础,或引发学术界广泛讨论与政策变革。
原文摘要 · Abstract (English)
The rapid advancement in building large language models (LLMs) has intensified competition among big-tech companies and AI startups. In this regard, model evaluations are critical for product and investment-related decision-making. While open evaluation sets like MMLU initially drove progress, concerns around data contamination and data bias have constantly questioned their reliability. As a result, it has led to the rise of private data curators who have begun conducting hidden evaluations with high-quality self-curated test prompts and their own expert annotators. In this paper, we argue that despite potential advantages in addressing contamination issues, private evaluations introduce inadvertent financial and evaluation risks. In particular, the key concerns include the potential conflict of interest arising from private data curators' business relationships with their clients (leading LLM firms). In addition, we highlight that the subjective preferences of private expert annotators will lead to inherent evaluation bias towards the models trained with the private curators' data. Overall, this paper lays the foundation for studying the risks of private evaluations that can lead to wide-ranging community discussions and policy changes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。