用泛化能力取代抽象智能,更可靠地评估大模型真实水平。
On the Measure of a Model: From Intelligence to Generality
- 以多任务学习视角重新定义评估,聚焦模型在新任务上的表现广度。
- 实证表明,只有泛化性经得起理论和实践检验,其他标准不可靠。
- 适合关注模型实际应用能力的研究者与开发者参考。
ARC、类瑞文测试和Blackbird任务等基准被广泛用于评估大语言模型的智能水平。然而,智能概念本身模糊且缺乏稳定定义,难以预测模型在问答、摘要或编程等实际任务中的表现。单纯优化这些基准可能导致评估目标与真实应用场景脱节。本文认为,评估应以泛化性为基础,而非抽象的智能概念。通过概念与形式分析,我们识别出智能评估常隐含的三个假设:泛化性、稳定性与真实性。研究发现,唯有泛化性能经受住理论和实证检验。智能并非促成泛化性的原因;相反,泛化性可视为多任务学习问题,直接关联评估结果的表现广度与可靠性。这一视角重新定义了人工智能进展的衡量方式,主张以泛化性作为跨多样化、持续演进任务的能力评估基础。
原文摘要 · Abstract (English)
Benchmarks such as ARC, Raven-inspired tests, and the Blackbird Task are widely used to evaluate the intelligence of large language models (LLMs). Yet, the concept of intelligence remains elusive- lacking a stable definition and failing to predict performance on practical tasks such as question answering, summarization, or coding. Optimizing for such benchmarks risks misaligning evaluation with real-world utility. Our perspective is that evaluation should be grounded in generality rather than abstract notions of intelligence. We identify three assumptions that often underpin intelligence-focused evaluation: generality, stability, and realism. Through conceptual and formal analysis, we show that only generality withstands conceptual and empirical scrutiny. Intelligence is not what enables generality; generality is best understood as a multitask learning problem that directly links evaluation to measurable performance breadth and reliability. This perspective reframes how progress in AI should be assessed and proposes generality as a more stable foundation for evaluating capability across diverse and evolving tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。