arXiv:2606.01434cs.CL2026-06

针对药物问答的虚假信息问题,提出基于权威源的智能代理与评测基准。

DrugClaw and DrugAudit: A Primary-Source-Grounded Agent and Authority-Aware Benchmark for Drug-Information Question Answering

论文配图:DrugClaw and DrugAudit: A Primary-Source-Grounded Agent and Authority-Aware Benchmark for Drug-Information Question Answering
图 1 · 摘自论文原文
  • 构建多智能体系统,通过反射式工作流检索原始药监数据
  • 在3772项评测中实现91.8%的权威源引用率,优于次优模型10.1个百分点
  • 提供可信赖的评测标准,适合医疗AI验证与临床决策支持研究

药物信息问答属于高风险场景,幻觉信息可能误导临床决策,且每个引用事实的来源可靠性同样重要。我们提出DrugClaw,一种基于多智能体的检索增强系统,通过反射驱动的状态机工作流查询药物与药警技能注册库,并返回基于原始监管或同行评审文献的答案。同时,我们构建了DrugAudit,一个包含3,772个样本的权威感知评测集,采用双评委大模型评分协议,评估上游来源匹配度、语义片段重叠及引用忠实性,评委间一致性kappa=0.88(几乎完美)。在DrugAudit及MedQA(751项)、PubMedQA(512项)的药物相关子集上,DrugClaw在各项指标中均排名第一:复合证据指数(双评委)、经评委评估的答案正确率、原始来源引用率(0.918,较次优提升10.1个百分点)、忠实性(0.887,提升5.9个百分点)、MedQA(0.920)、PubMedQA(0.693)。

原文摘要 · Abstract (English)

Drug-information question answering is a high-stakes setting where hallucinated facts can mislead clinical decision-making and the provenance of each cited fact matters as much as the fact itself. We present DrugClaw, a multi-agent retrieval-augmented system that queries a registry of drug and pharmacovigilance skills via a reflection-driven state-machine workflow and returns answers grounded in primary regulatory or peer-reviewed records. We also contribute DrugAudit, a 3,772-item authority-aware benchmark with an evaluation panel that scores upstream-of-gold source match, token-level semantic snippet overlap, and citation faithfulness under a dual-judge LLM-as-judge protocol with inter-judge kappa = 0.88 (almost-perfect). Across DrugAudit plus drug-related subsets of MedQA (751) and PubMedQA (512), DrugClaw is top-1 on every column of the headline table: composite Evidence Index under both judges, judge-mediated answer correctness, primary-source rate (0.918, +10.1 pp over next-best), faithfulness (0.887, +5.9 pp), MedQA (0.920), and PubMedQA (0.693).

药物问答可信AI评测基准多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。