用交叉熵游戏量化大模型的通用能力,突破生成文本局限。
Cross-Entropy Games for Language Models: From Implicit Knowledge to General Capability Measures
- 将大模型知识建模为概率分布,设计基于交叉熵的博弈任务
- 构建可扩展的博弈空间,涵盖摘要、反事实推理等10+能力场景
- 提出可迭代探索的演化式评估框架,适合评测模型通用智能
大语言模型(LLMs)在文本上定义了概率分布。通过探究其隐含知识的本质及算法意义,我们自然引出一系列超越生成采样的任务:包括摘要、反事实推理、异常检测、原创性搜索、逆向提示、辩论、创造性求解等。这些任务可被形式化为基于模型概率度量的博弈,称为交叉熵(Xent)游戏。Xent游戏可为单人或多人,涉及交叉熵得分与约束,以简单计算图和程序表达。我们证明该游戏空间足够丰富,能容纳大量有趣实例,且可由基本博弈论一致性公理构造。进一步讨论如何利用该空间衡量模型能力,提出构建Xent游戏度量:从给定范围中提取覆盖度量后生成的一组有限游戏,作为能力基准。为应对通用能力评估中的无限范围难题,我们提出采用受演化动力学启发的系统化探索策略。
原文摘要 · Abstract (English)
Large Language Models (LLMs) define probability measures on text. By considering the implicit knowledge question of what it means for an LLM to know such a measure and what it entails algorithmically, we are naturally led to formulate a series of tasks that go beyond generative sampling, involving forms of summarization, counterfactual thinking, anomaly detection, originality search, reverse prompting, debating, creative solving, etc. These tasks can be formulated as games based on LLM measures, which we call Cross-Entropy (Xent) Games. Xent Games can be single-player or multi-player. They involve cross-entropy scores and cross-entropy constraints, and can be expressed as simple computational graphs and programs. We show the Xent Game space is large enough to contain a wealth of interesting examples, while being constructible from basic game-theoretic consistency axioms. We then discuss how the Xent Game space can be used to measure the abilities of LLMs. This leads to the construction of Xent Game measures: finite families of Xent Games that can be used as capability benchmarks, built from a given scope, by extracting a covering measure. To address the unbounded scope problem associated with the challenge of measuring general abilities, we propose to explore the space of Xent Games in a coherent fashion, using ideas inspired by evolutionary dynamics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。