arXiv:2507.05266cs.CLcs.AI2025-07ACL被引 2

用用户行为预测评估大模型泛化能力,更可靠且低成本。

User Behavior Prediction as a Generic, Robust, Scalable, and Low-Cost Evaluation Strategy for Estimating Generalization in LLMs

  • 用用户行为预测代替传统任务衡量大模型泛化性。
  • GPT-4o在推荐任务上表现最优,但所有模型仍有提升空间。
  • 方法适合大规模、低资源场景下的模型评估。

衡量大语言模型(LLMs)的泛化能力面临数据泄露难题。随着模型规模扩大和计算成本下降,确保测试任务在训练中未出现几乎不可能实现。我们认为知识检索与推理任务不适合作为泛化评估指标,因模型并非针对特定任务训练。本文提出用户行为预测作为理论合理、可扩展且鲁棒的替代方案。我们构建新评估框架,在电影与音乐推荐数据集上测试 GPT-4o、GPT-4o-mini 与 Llama-3.1-8B-Instruct。结果符合框架预期:GPT-4o 表现优于 GPT-4o-mini 和 Llama,但所有模型均存在显著提升空间,尤其是 Llama。

原文摘要 · Abstract (English)

Measuring the generalization ability of Large Language Models (LLMs) is challenging due to data contamination. As models grow and computation becomes cheaper, ensuring tasks and test cases are unseen during training phases will become nearly impossible. We argue that knowledge-retrieval and reasoning tasks are not ideal for measuring generalization, as LLMs are not trained for specific tasks. Instead, we propose user behavior prediction, also a key aspect of personalization, as a theoretically sound, scalable, and robust alternative. We introduce a novel framework for this approach and test it on movie and music recommendation datasets for GPT-4o, GPT-4o-mini, and Llama-3.1-8B-Instruct. Results align with our framework's predictions, showing GPT-4o outperforms GPT-4o-mini and Llama, though all models have much room for improvement, especially Llama.

大模型评估用户行为预测泛化能力推荐系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。