arXiv:2602.15019cs.AIcs.IR2026-02被引 2

用自学习智能体在全球非英文渠道中精准发现潜在药物资产,避免投资损失。

Hunt Globally: Wide Search AI Agents for Drug Asset Scouting in Investing, Business Development, and Competitive Intelligence

  • 构建树状结构自学习代理,跨多语言源进行无幻觉搜索。
  • 在多语言基准测试中达79.7%的F1分数,显著优于主流模型。
  • 适合投资、生物医药开发和竞争情报领域需全面覆盖的团队使用。

生物制药创新重心已转移:多数新药资产源于美国以外地区,且主要通过区域性非英语渠道披露。数据显示,超85%的专利申请来自美国以外,其中中国占全球近半;中国也贡献了全球30%的药物研发,涵盖1200多个候选药物。在此高风险环境下,遗漏“未被关注”的资产可能导致数亿美元损失,资产搜寻成为决定性的竞争环节。然而当前深度研究型AI代理在异构、多语言数据中仍难以实现高召回率且不产生幻觉。本文提出一种药物资产搜寻基准方法,并设计一个经过调优的树状结构自学习生物眼智体(Bioptic Agent),以实现完整、无幻觉的搜寻。构建了一个挑战性完整性基准,包含多语言多代理流水线:复杂用户查询与真实目标资产匹配,且这些资产大多不在美国中心视角下。为反映真实交易复杂度,收集了专家投资者、业务发展与风投人员的筛选查询作为先验,用于条件生成基准查询。评估采用经专家意见校准的LLM评判体系。在该基准上,本模型取得79.7% F1分数,优于Gemini 3.1 Deep Think(59.2%)、Gemini 3.1 Pro Deep Research(58.6%)、Claude Opus 4.6(56.2%)、OpenAI GPT-5.2 Pro(46.6%)、Perplexity Deep Research(44.2%)及Exa Websets(26.9%)。性能随计算资源增加而显著提升,验证了更多算力带来更优结果的观点。

原文摘要 · Abstract (English)

Bio-pharmaceutical innovation has shifted: many new drug assets now originate outside the United States and are disclosed primarily via regional, non-English channels. Recent data suggests over 85% of patent filings originate outside the U.S., with China accounting for nearly half of the global total; a growing share of scholarly output is also non-U.S. Industry estimates put China at 30% of global drug development, spanning 1,200+ novel candidates. In this high-stakes environment, failing to surface "under-the-radar" assets creates multi-billion-dollar risk for investors and business development teams, making asset scouting a coverage-critical competition where speed and completeness drive value. Yet today's Deep Research AI agents still lag human experts in achieving high-recall discovery across heterogeneous, multilingual sources without hallucinations. We propose a benchmarking methodology for drug asset scouting and a tuned, tree-based self-learning Bioptic Agent aimed at complete, non-hallucinated scouting. We construct a challenging completeness benchmark using a multilingual multi-agent pipeline: complex user queries paired with ground-truth assets that are largely outside U.S.-centric radar. To reflect real deal complexity, we collected screening queries from expert investors, BD, and VC professionals, and used them as priors to conditionally generate benchmark queries. For grading, we use LLM-as-judge evaluation calibrated to expert opinions. On this benchmark, our Bioptic Agent achieves 79.7% F1 score, outperforming Gemini 3.1 Deep Think (59.2%), Gemini 3.1 Pro Deep Research (58.6%), Claude Opus 4.6 (56.2%), OpenAI GPT-5.2 Pro (46.6%), Perplexity Deep Research (44.2%), and Exa Websets (26.9%). Performance improves steeply with additional compute, supporting the view that more compute yields better results.

药物发现AI agent多语言检索投资情报

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。