构建可控诊断基准,测试智能体发现隐藏结构的能力。
Auto-Discovery-Bench: Diagnosing Structured State Tracking in Oracle-Guided Discovery
- 通过假设-干预-反馈循环,测试智能体在确定性环境中的结构发现能力。
- 变量数、轨迹长度和干扰项增加时,模型性能显著下降。
- 揭示长期结构信息维护是交互式科学探索的关键瓶颈。
交互式发现要求智能体在多轮反馈中持续维护和更新结构化信念。在真实、嘈杂的科学环境中评估前,有必要在受控条件下隔离这一核心能力。我们提出 Auto-Discovery-Bench,一个基于确定性预言者引导的诊断基准,智能体通过重复的假设-干预-反馈循环来恢复隐藏结构。该基准涵盖三种受控发现抽象:有向图发现、无向关系发现和符号方程发现。实验显示,随着变量数量、轨迹长度和干扰项的增加,各模型性能均下降。额外的轨迹追踪诊断表明,即使移除干预选择和假设生成,许多失败仍持续存在,说明维持和整合长程结构信息是预言者引导发现的重要瓶颈。Auto-Discovery-Bench并非替代真实发现环境,而是提供可复现、低混淆的诊断平台,用于分离交互式科学智能体所需的基础能力。
原文摘要 · Abstract (English)
Interactive discovery requires agents to maintain and update structured beliefs over many rounds of feedback. Before evaluating agents in noisy, open-ended scientific environments, it is useful to isolate this prerequisite capability under controlled conditions. We introduce Auto-Discovery-Bench, a deterministic oracle-guided diagnostic benchmark in which agents recover hidden structures through repeated hypothesis--intervention--feedback cycles. The benchmark instantiates three controlled discovery abstractions: directed graph discovery, undirected relational discovery, and symbolic equation discovery. Across models, performance degrades as the number of variables, trajectory length, and distractors increase. A separate trajectory-tracking diagnostic shows that many failures persist even when intervention selection and hypothesis generation are removed, suggesting that limitations in maintaining and integrating long-range structured information are an important bottleneck for oracle-guided discovery. Auto-Discovery-Bench is not intended to replace realistic discovery environments; rather, it provides a reproducible, low-confound diagnostic testbed for isolating a prerequisite capability for interactive scientific agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。