arXiv:2503.21138cs.AIcs.LG2025-03

用计算理论加速小智能体评估,确保因果可靠性。

A Computational Theory for Efficient Mini Agent Evaluation with Causal Guarantees

  • 构建评估模型替代实验,通过元学习处理异构智能体空间。
  • 12个场景下评估误差降低24.1%至99.0%,耗时减少3到7个数量级。
  • 适合需要高效、可信评估的智能体研发与部署场景。

为降低智能体实验评估成本,本文提出一种面向小型智能体的计算评估理论:构建评估模型以加速评估过程。我们证明了无限智能体情况下评估模型的泛化误差和因果效应误差上界,并验证其估计因果效应的效率与一致性。针对异构智能体空间问题,提出一种元学习方法来训练评估模型。相比现有方法,本方法在个体医疗、科学模拟、社会实验、商业活动及量子交易等12个场景中,评估误差减少24.1%至99.0%;每次评估耗时相比实验或仿真减少3至7个数量级。

原文摘要 · Abstract (English)

In order to reduce the cost of experimental evaluation for agents, we introduce a computational theory of evaluation for mini agents: build evaluation model to accelerate the evaluation procedures. We prove upper bounds of generalized error and generalized causal effect error of given evaluation models for infinite agents. We also prove efficiency, and consistency to estimated causal effect from deployed agents to evaluation metric by prediction. To learn evaluation models, we propose a meta-learner to handle heterogeneous agents space problem. Comparing with existed evaluation approaches, our (conditional) evaluation model reduced 24.1\% to 99.0\% evaluation errors across 12 scenes, including individual medicine, scientific simulation, social experiment, business activity, and quantum trade. The evaluation time is reduced 3 to 7 order of magnitude per subject comparing with experiments or simulations.

智能体评估因果推断元学习高效评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。