arXiv:2607.05155cs.CLcs.LG2026-07被引 3

发现真实环境学习性能遵循高精度对数正弦缩放规律。

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments

论文配图:EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments
图 1 · 摘自论文原文
  • 基于3.8万小时真实交互数据,发现学习性能随时间呈对数正弦增长。
  • 模型学习速度每三个月约翻倍,跨代际提升显著。
  • 公开134项真实任务与评估框架,助力智能体持续学习研究。

预训练缩放定律表明模型能力随数据和算力提升而可预测增长,但部署后从真实环境学习仍不明确。分析跨134个真实任务、约3.8万小时的智能体环境交互数据,首次发现学习过程中的整体性能遵循对数正弦缩放规律,拟合度极高(R² = 0.998)。跨模型代际观察到,智能体学习速度每三个月大致翻倍。该发现源自EdgeBench——一套包含134个真实世界任务的评测集,涵盖科学发现、软件工程、组合优化、专业知识工作、形式数学及互动游戏,每个任务需连续运行至少12小时并接受多层次反馈,均由专家投入大量精力构建。我们公开发布其中51项任务及完整评估框架,以加速对智能体从真实经验中学习机制的研究。

原文摘要 · Abstract (English)

Pretraining scaling laws reveal that model capability improves predictably with data and compute. But learning from real world environments after deployment remains far less understood. Analyzing roughly 38,000 hours of agent interaction with the environment across 134 real world tasks, we find, to the best of our knowledge, the first evidence that overall performance during environment learning follows a log-sigmoid scaling law with remarkably high precision, reaching R^2 = 0.998. Across model generations, we also find that agent learning speed roughly doubles every three months. This discovery stems from EdgeBench, a suite of 134 real world tasks with ultra-long horizons, spanning scientific discovery, software engineering, combinatorial optimization, professional knowledge work, formal mathematics, and interactive games. Each task sustains at least 12 hours of continuous agent operation under rich, multilevel feedback, and is built through substantial expert effort. We publicly release 51 tasks and our full evaluation framework to accelerate the study of how agents learn from real world experience.

智能体学习缩放定律真实环境评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。