arXiv:2608.11341cs.AI2026-08

构建可验证的发现型AI评估框架,推动AI从解题走向真实世界创新。

Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence

  • 用系统化流程筛选20个高价值现实问题,构建可追踪的探究环境。
  • 在病毒衣壳设计中提升7%性能,在药物重定位任务中最高提升7.6分。
  • 适合追求真实科学发现的科研团队与具身智能研发者使用。

当前前沿模型虽能解决明确任务,但面对真实世界复杂挑战时缺乏可执行性与可验证性。本文提出Apodex Discovery框架,通过重型求解器(heavy-duty solver)实现长期、状态持续、可验证的探究。该系统包含三大核心:一是覆盖16个行业561项产业的难题普查,筛选出423个高价值问题,首批发布20个;二是统一的环境-任务-剧集抽象结构,支持数据、工具、约束、反馈、轨迹记录及中间成果验证;三是HDS6评估体系独立衡量工具、修复、替代、连贯性、证据与范围等维度。在AAV衣壳设计中,其性能超越现有最佳方案7%;在药物重定位与再配方任务中,基于生物医学环境的定制系统使GPT-5.5和GPT-5.6-sol的平均归一化预测得分分别提升2.5与7.6分。受控消融实验表明,固定TRACES剧集接口可精准归因性能差异至具体组件。该框架将AI评估从预设基准推进至面向真实发现的可验证探究。

原文摘要 · Abstract (English)

Apollo did not reach the Moon merely because its engineers could solve difficult equations. It succeeded by turning a distant ambition into a mission architecture of explicit objectives, simulation, verification, and repeated correction. AI now faces a similar transition: frontier models can solve difficult tasks once the problem, tools, and success criteria are specified, yet consequential real-world challenges rarely arrive in an executable or verifiable form. We introduce Apodex Discovery, a framework for building and evaluating discoverative AI through the heavy-duty solver, a system comprising a foundation model, harness, tools, and control policies that pursues extended, stateful, verifiable investigations. It has three core components. First, a problem-scouting process surveyed 561 industries across 16 sectors, assembled 423 high-value real-world problems, and selected 20 for the initial release. Second, a common environment-task-episode abstraction provides data, tools, constraints, feedback, trajectory recording, and verification of intermediate artifacts and final submissions. Third, HDS6 evaluates Tools, Repair, Alternatives, Coherence, Evidence, and Scope independently of final-task success. In AAV capsid design, Apodex surpassed the published state of the art by 7% across viability, tropism, structure prediction, and generative design. In drug repurposing and reformulation, a task-specific biomedical environment improved the mean normalized prediction score of GPT-5.5 and GPT-5.6-sol by 2.5 and 7.6 points over the same closed-book backbone. Controlled ablations show that the fixed TRACES episode interface enables attribution of performance differences to specific solver components. Apodex Discovery moves AI evaluation beyond predefined benchmarks toward verifiable investigations aimed at genuine discovery.

发现型AI评估框架科学发现可验证性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。