提出APTBench,评估大模型预训练阶段的智能体潜力。
APTBench: Benchmarking Agentic Potential of Base LLMs During Pre-Training
- 将真实任务转化为多选题/填空题,适配基础模型
- 聚焦规划与执行能力,覆盖软件工程等场景
- 轻量高效,比后训练评估更早预测智能体表现
随着基于大语言模型的智能体快速发展,将智能体相关数据融入大模型预训练阶段成为趋势,以更好对齐模型与真实世界自主任务执行。然而,现有预训练评估主要关注孤立静态技能(如常识、数学或代码推理),无法反映模型的智能体能力;而现有智能体评估通常针对微调后的模型,需多轮任务执行能力,基础模型难以支撑。为此,我们提出APTBench,将真实世界智能体任务及成功轨迹转化为适配基础模型的多选题或文本补全题,聚焦规划与行动等核心智能体能力,涵盖软件工程与深度研究等关键场景。相比通用基准,APTBench能更有效预测模型作为智能体的下游表现,同时显著低于完整端到端后训练评估的成本与开销。
原文摘要 · Abstract (English)
With the rapid development of LLM-based agents, there is a growing trend to incorporate agent-specific data into the pre-training stage of LLMs, aiming to better align LLMs with real-world autonomous task execution. However, current pre-training benchmarks primarily focus on isolated and static skills, e.g., common knowledge or mathematical/code reasoning, and fail to reflect model's agentic capabilities. On the other hand, agent benchmarks are typically designed for post-trained models, requiring multi-turn task execution abilities that base models struggle to support. Thus, there is a compelling need for a benchmark that can evaluate agentic potentials during pre-training and guide the model training more effectively. To address this gap, we propose APTBench, a framework that converts real-world agent tasks and successful trajectories into multiple-choice or text completion questions tailored for base models. It focuses on core agentic abilities, e.g., planning and action, and covers key agent scenarios, software engineering and deep research. Compared to existing general-purpose benchmarks, APTBench offers a more predictive signal of a model's downstream performance as an agent, while remaining significantly more lightweight and cost-effective than full-scale, end-to-end agent evaluations after post-training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。