arXiv:2512.03065cs.LGcs.AI2025-12

用强化学习让生命科学AI agents自动优化决策,不依赖标注数据。

Optimizing Life Sciences Agents in Real-Time using Reinforcement Learning

  • 基于用户反馈,用贝叶斯上下文猜拳算法动态选择生成策略、工具和领域路由。
  • 在真实查询上提升15%-30%用户满意度,20-30轮后学习效果显著。
  • 无需标签数据,可持续适应用户偏好,适合需要灵活响应的科研助手场景。

生命科学领域的生成式AI代理面临核心挑战:如何为从简单事实到复杂机制推理的多样化问题选择最优处理方式。传统方法依赖固定规则或昂贵的标注数据,难以适应变化的条件或用户偏好。本文提出一种新框架,将AWS Strands Agent与汤普森采样上下文赌博机结合,仅通过用户反馈即可让AI代理学习最优决策策略。系统优化三个关键维度:生成策略选择(直接生成 vs. 思维链)、工具选择(文献检索、药物数据库等)以及领域路由(药理学、分子生物学、临床专家)。在真实生命科学查询上的实证评估显示,相比随机基线,用户满意度提升15%-30%,且在20-30次交互后出现清晰的学习模式。该方法无需真实标签,能持续适应用户偏好,为智能体系统中的探索-利用权衡提供了严谨解决方案。

原文摘要 · Abstract (English)

Generative AI agents in life sciences face a critical challenge: determining the optimal approach for diverse queries ranging from simple factoid questions to complex mechanistic reasoning. Traditional methods rely on fixed rules or expensive labeled training data, neither of which adapts to changing conditions or user preferences. We present a novel framework that combines AWS Strands Agents with Thompson Sampling contextual bandits to enable AI agents to learn optimal decision-making strategies from user feedback alone. Our system optimizes three key dimensions: generation strategy selection (direct vs. chain-of-thought), tool selection (literature search, drug databases, etc.), and domain routing (pharmacology, molecular biology, clinical specialists). Through empirical evaluation on life science queries, we demonstrate 15-30\% improvement in user satisfaction compared to random baselines, with clear learning patterns emerging after 20-30 queries. Our approach requires no ground truth labels, adapts continuously to user preferences, and provides a principled solution to the exploration-exploitation dilemma in agentic AI systems.

AI代理强化学习生命科学自适应系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。