arXiv:2606.09556cs.AI2026-06

AI制药评估的上限由数据质量决定,而非推理能力。

AI Scientists Are Only as Good as Their Evidence: A Stratified Ablation of Proprietary Data and Reasoning Skills in Drug-Asset Valuation

论文配图:AI Scientists Are Only as Good as Their Evidence: A Stratified Ablation of Proprietary Data and Reasoning Skills in Drug-Asset Valuation
图 1 · 摘自论文原文
  • 对比三组实验:纯网络模型、公开工具+规则框架、加入专有药物数据集
  • 专有数据使决策准确率从0.38提升至0.96,长尾资产达0.93
  • 即使推理完美,无数据覆盖仍限决策上限至3.83

AI科学家在药物资产估值中常被评估为模型质量或推理结构决定性能。我们提出新假设:知识密集型科学决策的关键限制在于可访问的证据基础。在13个资产分层基准上进行三组对照实验:A为仅用网络的LLM分析员,B增加公开结构化工具与14维估值手册、验证器、客观性策略及红队机制,C在此基础上加入专有诺亚AI药物管线、试验与交易情报数据集。结果显示,B显著提升校准度与审计纪律(区间准确率从0.80升至0.89,客观性从3.16升至3.30),但未突破事实天花板。在能力超集统计下,A和B仅恢复0.25与0.38的黄金基准表现,而C恢复0.96;在专有长尾子集上,C达0.93,远超A/B的0.26/0.30。盲评决策质量相近(7.01 vs 6.96),故引入‘完备性感知决策效用’:决策质量×黄金覆盖率。此时C达7.43,远超A/B的1.76/2.57。即便非专有报告完美,受限于覆盖范围,上限仍为3.83。结论并非推理不重要,而是专有证据设定了认知与决策的上限。

原文摘要 · Abstract (English)

AI Scientist agents are often evaluated as if capability were mainly a function of model quality, prompting, or reasoning scaffolds. We test a different hypothesis in drug-asset valuation: for knowledge-intensive scientific decisions, the limiting factor is often the evidence substrate the agent can access. We run a controlled three-arm ablation on a production valuation agent: A is a plain web-only LLM analyst, B adds public structured tools plus a 14-dimension valuation playbook, verifier, objectivity policy and red-team, and C adds the proprietary Noah AI corpus of curated pipeline, trial and deal intelligence. Across a 13-asset stratified benchmark, B improves calibration and audit discipline: tier-in-range accuracy rises from 0.80 to 0.89 and objectivity from 3.16 to 3.30. But B does not remove the factual ceiling. Under capability-superset accounting, A and B recover only 0.25 and 0.38 of the curated gold competitive record, while C recovers 0.96; on the curated long-tail subset, C reaches 0.93 vs. 0.26/0.30. Raw blind-panel decision quality is similar for A and B (7.01 vs. 6.96), so we introduce completeness-aware decision utility: informed decision-quality = decision-quality x gold-coverage. On this metric, C reaches 7.43 vs. 1.76/2.57 for A/B. Even a perfect non-proprietary-data report would be capped at 3.83 by B's coverage. The result is not that reasoning scaffolds are unimportant; they improve calibration and discipline. Rather, proprietary evidence sets the upper bound of what the AI Scientist can know and therefore decide.

AI制药证据依赖数据壁垒决策效用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。