arXiv:2608.25215cs.AIcs.MA2026-08

用低成本确定性策略可实现接近顶尖的蛋白质功能预测,适合常规任务。

Federation Is Nearly Free, Reasoning Is Not: Tradeoffs for AI Co-Scientists in Protein Characterization Workflows

  • 采用强化学习策略替代大模型推理,零成本达88%准确率
  • 顶级大模型(Opus)准确率达92%-94%,但成本高且不可复现
  • 联邦架构几乎不影响性能,适合跨机构协作场景

自然语言驱动的自主科学代理工作流在灵活性与推理能力之间存在根本权衡,牺牲了确定性、可复现性和可观测性。本文在真实科学代理平台中系统评估了这一权衡,任务为根据蛋白序列判断其功能。对比了联邦拓扑、经典强化学习(PPO)与大模型驱动方法、不同语言模型及提示工程效果,并按蛋白新颖性分层分析。结果发现,大模型选择对预测质量影响远超拓扑或提示(Opus达92%-94%,o4-mini仅40%-50%)。PPO策略在零令牌成本下达到88%准确率,延迟最低且完全一致,但无推理过程。专家提示的大模型精度最高,但成本高且一致性差;难度越大,提示依赖越强。联邦架构对性能影响可忽略。研究建议:常规可验证任务宜用低成本确定性策略,开放探索则保留灵活大模型推理。

原文摘要 · Abstract (English)

Natural language driven autonomous co-scientist workflows involve a fundamental trade-off between flexibility and reasoning at the expense of determinism, reproducibility, and observability. Such agents increasingly must communicate across institutional boundaries, where federation topology can shape latency and cost. We systematically evaluated these tradeoffs using a controlled ablation on a production agentic platform for science. We use a verifiable task: given a protein sequence, we ask an agent to confidently characterize its function by routing across common tools. We compare federation topology, classic RL vs LLM-driven harnesses, language model, and prompt expertise. We also stratify results by protein novelty. We find that the choice of LLM dominated prediction quality far more than topology or prompting (Opus ~92%-94% vs o4-mini ~40%-50%). The PPO policy was nearly as accurate as the best LLM (88%) at zero token cost, fastest latency, and perfect consistency, but yields no reasoning trace. Expert prompted LLMs reached the highest accuracy but were high-cost and less consistent; prompt dependence was largest when the task was hardest. Federation imposed a negligible penalty on performance. These results offer actionable guidance for deploying agents for scientific workflows: for routine, verifiable tasks, a cheap deterministic policy delivers near-frontier accuracy with complete reproducibility, while flexible LLM reasoning is best reserved for open-ended discovery.

科学智能大模型联邦学习蛋白质

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。