arXiv:2608.21057cs.LG2026-08

构建可信赖的智能药研评估系统,让AI评委与专家意见对齐。

Designing a Robust LLM-Based Evaluation System for Agentic AI in Drug Discovery Through Human Alignment

论文配图:Designing a Robust LLM-Based Evaluation System for Agentic AI in Drug Discovery Through Human Alignment
图 1 · 摘自论文原文
  • 定义四大评价维度+工具调用校验,构建多维评估体系。
  • 通过专家对比验证,优化后裁判与人类共识一致率达0.86。
  • 发现非正式表达反而有助于提升输出质量,支持重写问题再提问。

智能体大模型正重塑化学与药物研发流程,但其开放性、工具增强型输出的评估仍是核心瓶颈。传统基于参考文本的指标(如BLEU、ROUGE)无法捕捉语义正确性,而人工评估又难以匹配系统迭代速度。现有‘以LLM为裁判’的方案缺乏与专家意见对齐的验证。本文针对阿斯利康部署的ChatInvent智能药研助手,提出一个可对齐人类专家的LLM评估框架,贡献如下:首先,定义了完整性、相关性、结构清晰度、范围符合性四维评价标准,并加入确定性工具调用正确性检查;其次,通过五名专家标注对比,验证Gemini 3.1 Pro、Claude Opus 4.7、GPT-5和Llama 3.1 70B等候选裁判的对齐效果;第三,采用少量人类标注示例进行少样本提示优化,使最优裁判与人类多数意见的一致性从0.80提升至0.86;第四,应用于70个保留问题,揭示系统局限性,发现非正式表述不会系统性降低输出质量,反而建议在查询前让LLM重写原始问题。该框架为科学领域智能体系统的可信评估提供可复用模板。

原文摘要 · Abstract (English)

Agentic large language model (LLM) systems are reshaping scientific workflows in chemistry and drug discovery, but evaluating their open-ended, tool-augmented outputs remains a fundamental bottleneck. Reference-based metrics such as BLEU and ROUGE fail to capture semantic correctness, while expert human evaluation does not scale to the iteration speed these systems demand. The LLM-as-a-Judge paradigm has emerged as a scalable alternative, but existing drug discovery benchmarks deploy LLM judges without validating their alignment with human experts. In this work, we present an LLM-as-a-Judge evaluation framework for ChatInvent, an agentic drug discovery assistant deployed at AstraZeneca, with four contributions. First, we define four output-quality evaluation dimensions---Completeness, Relevancy, Structural Clarity, and Scope Adherence---alongside deterministic Tool Call Correctness checks. Second, we validate the judge through a human alignment study with five expert annotators, comparing Gemini 3.1 Pro, Claude Opus 4.7, GPT-5, and Llama 3.1 70B as candidate judges. Third, we optimize the best-performing judge using few-shot demonstrations of human-annotated examples, improving alignment with the human majority vote from 0.80 to 0.86. Fourth, applying the optimized judge to 70 held-out questions, we surface concrete limitations and find that informal phrasings do not systematically degrade output quality; if anything, it is helpful to have the LLM rewrite the original question before querying the agent. Our framework provides a reusable template for human-aligned evaluation of agentic systems in scientific domains.

AI制药评估框架大模型评测人机对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。