用实验设计方法量化大模型自主建模的变异性与效率
An Experimental Design Approach to Evaluating Agentic AI's Autonomous Model Discovery

- 将智能体视为随机建模算子,控制推理力度等变量系统评估其表现
- 发现推理投入与建模成本、过程复杂度存在非线性关系,存在最优平衡点
- 适合研究智能体自主性、自动化建模效率的科研人员和工程团队
大语言模型编程智能体在开放式数据建模与分析中日益活跃。由于其具有随机性和自适应性,单次基准测试难以充分刻画其自主建模行为。本文提出一种实验设计与分析框架,系统评估该发现过程,量化其变异性并识别关键影响因素。将智能体视为随机建模算子,将任务特定数据与优化目标映射为拟合模型。重点考察Codex与Claude Code两个算子,在推理努力、任务类型、优化指标及训练数据构成等受控条件下,对输出质量、成本、耗时与过程复杂度等多响应进行回归分析。进一步提出效用对齐的规范分解,以刻画推理努力的主导方向,并评估其是否与性能-成本效用方向一致。框架在联网词形生成游戏测试集上验证,揭示了推理投入与成本、过程复杂度之间的深层关系。
原文摘要 · Abstract (English)
Large language model coding agents increasingly perform open-ended data modeling and analysis. These agents are stochastic and adaptive, and therefore their autonomous model discovery behavior cannot be adequately characterized by a single benchmark run. In this work, we propose an experimental design and analysis framework for systematically evaluating this discovery process, quantifying its variability, and identifying important factors. The proposed framework treats these agents as stochastic model-discovery operators, which map task-specific discovery data and an optimization target to a fitted model. Specifically, we investigate two such operators, Codex and Claude Code, under controlled experimental factors including agent's reasoning effort, task, optimization metric, and composition of training data. For each agent-task-metric combination, regression models and inference are conducted for multiple responses such as output quality, dollar cost, wall-clock time, and process complexity. Furthermore, we develop a utility-aligned canonical decomposition to characterize the dominant direction of the reasoning-effort effect and to assess whether that direction aligns with a performance-cost utility direction. The proposed framework is demonstrated on a testbed of networked word-forming games with insightful findings on reasoning effort with respect to cost and process complexity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。