提出决策评估框架,让AI自动机器学习系统更透明可靠。
A Framework for Assessing AI Agent Decisions and Outcomes in AutoML Pipelines
- 设计观察型评估代理,从四个维度审查中间决策
- 可检测91.9%的错误决策,识别出与最终结果无关的逻辑矛盾
- 能追溯决策对结果的影响,范围达-4.9%至+8.3%
基于智能体的自动化机器学习系统依赖大语言模型在数据处理、模型选择和评估等多阶段做出复杂决策。然而现有评估方法仍以最终任务性能为中心,缺乏对中间决策质量的结构化评估。我们发现已有智能体式AutoML系统均未报告用于事后分析的决策级评估指标。为此,我们提出评估代理(Evaluation Agent, EA),作为无干扰的观察者,从决策有效性、推理一致性、模型质量风险(超越准确率)以及反事实决策影响四个维度评估中间决策。在四项概念验证实验中,EA能够以0.919的F1分数检测故障决策,独立识别推理不一致,且可归因于决策的下游性能变化,影响范围为-4.9%至+8.3%。这些结果表明,决策中心评估能揭示仅看最终结果时无法察觉的失败模式。本工作将智能体式AutoML系统的评估范式从结果导向转向决策审计,为构建可信赖、可解释、可监管的自主机器学习系统奠定基础。
原文摘要 · Abstract (English)
Agent-based AutoML systems rely on large language models to make complex, multi-stage decisions across data processing, model selection, and evaluation. However, existing evaluation practices remain outcome-centric, focusing primarily on final task performance. Through a review of prior work, we find that none of the surveyed agentic AutoML systems report structured, decision-level evaluation metrics intended for post-hoc assessment of intermediate decision quality. To address this limitation, we propose an Evaluation Agent (EA) that performs decision-centric assessment of AutoML agents without interfering with their execution. The EA is designed as an observer that evaluates intermediate decisions along four dimensions: decision validity, reasoning consistency, model quality risks beyond accuracy, and counterfactual decision impact. Across four proof-of-concept experiments, we demonstrate that the EA can (i) detect faulty decisions with an F1 score of 0.919, (ii) identify reasoning inconsistencies independent of final outcomes, and (iii) attribute downstream performance changes to agent decisions, revealing impacts ranging from -4.9% to +8.3% in final metrics. These results illustrate how decision-centric evaluation exposes failure modes that are invisible to outcome-only metrics. Our work reframes the evaluation of agentic AutoML systems from an outcome-based perspective to one that audits agent decisions, offering a foundation for reliable, interpretable, and governable autonomous ML systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。