提出新规划框架,解决科学发现中能力构建的盲区问题。
Capability-Gated Planning: Cost-to-Goal Discovery and the Limits of Myopic Experiment Selection
- 将科学发现建模为信念空间中的最短路径问题,引入能力改变动作图的新机制。
- 证明任何只看短期信息收益的规划器在特定场景下可能永远无法达成目标。
- 设计具备能力感知的启发式函数,显著提升长期规划能力,适合复杂探索任务。
自动化科学发现系统需反复决定执行何种实验、测试何种假设或构建何种工具,并判断何时停止。当前许多系统通过最大化短期指标(如单位成本预期信息增益或学习到的可信度分数)做决策。本文揭示该方法存在结构性缺陷:某些行动具有建设性——它们获取认知能力(如仪器、检测方法、流程、模拟器或抽象),其价值不在于即时信息,而在于未来可支持的动作。当达到可靠答案的最低成本路径需要一系列此类构建时,仅评估有限视野内信息收益的规划器无法重视首个构建,因其在视野内无直接信息输出,会被任意有微小信息回报的测量行动压制。本文将目标导向发现建模为信念空间中的随机最短路径问题,其中建设性实验会改变下游动作图,并证明对每个前瞻深度d,均存在实例使得所有仅追求短期信息最大化的规划器近似比无界,且另一实例中其根本无法到达目标。核心机制为‘能力不可辨识引理’:在视野内,获取能力与支付空操作在观测上无法区分。这确立了能力屏蔽作为独立于曲率(次模性)和信息顺序(自适应差距)的可达性难题维度。提出CG-Plan,一种具备能力感知代价到目标启发式(h = h_cap + h_exp)的增量重规划器。在受控测试平台上,性能差距仅在能力屏蔽条件下出现,且对固定前瞻深度持续存在,尤其当近似假设来自数据一致的提议者时显现。
原文摘要 · Abstract (English)
Systems that automate scientific discovery must repeatedly decide which experiment to run, which hypothesis to test, which tool to build, and when to stop. Many systems make these decisions by maximizing a myopic score such as expected information gain per unit cost or a learned plausibility score. We identify a structural limitation of this approach. Some actions are constructive: they acquire an epistemic capability (an instrument, assay, pipeline, simulator, or abstraction) whose value lies not in the information returned immediately but in the future actions it makes available. When the least-cost route to a confident answer requires a chain of such constructions, a planner that scores actions only by information obtainable within a bounded horizon cannot value the first construction: it yields no information within the horizon and is dominated by any measurement with positive information, however small. We formulate goal-directed discovery as a stochastic shortest-path problem in belief space in which constructive experiments change the downstream action graph, and prove that for every lookahead depth d there is an instance on which every myopic information-maximizing planner has an unbounded approximation ratio, and a related instance on which it never reaches the goal. The mechanism is a capability-indistinguishability lemma: within the horizon, acquiring a capability can be observationally indistinguishable from paying for a null action. This establishes capability gating as a reachability axis of difficulty distinct from curvature (submodularity) and information order (adaptivity gaps). We introduce CG-Plan, an incremental replanner with a capability-aware cost-to-go heuristic h = h_cap + h_exp. In a controlled testbed, the performance gap appears only under gating, persists for every fixed horizon, and arises when near-miss hypotheses come from a data-consistent proposer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。