arXiv:2607.18696cs.AI2026-07

用世界模型取代部门架构,让AI生物公司更高效决策。

Do AI-Native Biotechs Need Departments? Benchmarking Company World Models for AI-Driven Drug Development

  • 构建公司世界模型,统一管理资产与价值的动态演化。
  • 价值转化架构在模拟中得分最高,优于传统部门模式。
  • 适合研究AI驱动药物研发的组织设计与智能体系统。

AI原生生物技术公司常沿用人类生物公司的组织架构,将角色映射为代理。本文提出新抽象:公司世界模型,即一个持久的资产-价值状态表示,包含转移模型、显式价值函数、规划能力及跨科学、监管、商业拓展、财务和执行约束的持续更新。我们引入干实验基准,测试AI代理组织应模仿部门还是围绕该世界模型运作。基准包含45个基于公开信息的回顾性决策案例,设有严格时间限制、隐藏结果、统一模式、自动评分及盲评对战。对比了人类组织模仿型、强化版人类模仿型、以资产为中心的AI原生架构,以及以价值转换为核心的架构。后者通过交易、审批、收入与投资仲裁循环更新实时资产价值记录,在外部商业拓展、监管批准、上市及收入纪律的成功指标下,取得最高自动评分,并被价值导向的盲评专家显著偏好。压力测试显示,强人类基线仍具竞争力,中立评委未见显著优势。仅使用Codex的机制消融表明,收入室、交易室和审批室在目标下具备实际作用。核心发现是目标敏感性:部门可能仍是有效治理视角,但AI原生运营的核心应是共享的、可预测的资产-价值状态,而非静态的人类组织图。本研究为纯干实验,未验证真实药物成功、临床获益或收入预测准确性。

原文摘要 · Abstract (English)

AI-native biotechnology companies are often designed by copying human biotech org charts into agent roles. We argue for a different abstraction: a Company World Model, defined as a persistent asset-to-value state representation with transition models, explicit value functions, planning, and updating across scientific, regulatory, BD, commercial, financial, and execution constraints. We introduce a dry-lab benchmark for testing whether AI-agent organizations should mimic departments or operate around such a world model. The benchmark contains 45 retrospective public-information decision cases with strict time cutoffs, hidden outcomes, common schemas, automatic scoring, and blinded pairwise judging. We compare human-org-mimic, stronger human-org-mimic-plus, AI-native asset-centric, and AI-native value-conversion architectures. The value-conversion architecture is a prompt-level approximation of a Company World Model: a Live Asset Value Record updated by Deal, Approval, Revenue, and Investment Arbiter loops. Under a success function defined by external BD, regulatory approval and launch, and revenue discipline, it achieved the highest automatic value-conversion score and was strongly preferred over the original baselines by value-specific blinded judges. Stress tests narrowed the claim: a stronger human baseline remained competitive, and a neutral judge did not show robust value-conversion dominance. Codex-only mechanistic ablations suggest that Revenue Room, Deal Room, and Approval Room carry useful work under the target objective. The central finding is objective-sensitive: departments may remain useful governance views, but the core AI-native operating primitive should be a shared, predictive asset-to-value state rather than a static human org chart. The study is dry-lab only and does not establish real-world drug success, clinical benefit, or revenue prediction accuracy.

AI制药组织架构世界模型智能体系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。