arXiv:2604.20441cs.AI2026-04

为医疗研究智能体设计专用评估框架,提升其科学可靠性与部署安全性。

MedSkillAudit: A Domain-Specific Audit Framework for Medical Research Agent Skills

论文配图:MedSkillAudit: A Domain-Specific Audit Framework for Medical Research Agent Skills
图 1 · 摘自论文原文
  • 构建分层审计框架,系统化评估医疗研究技能的发布准备度。
  • 75项技能中近六成未达有限发布标准,系统评分一致性优于人工专家。
  • 针对协议设计类技能效果最好,写作类存在评分体系与专家认知偏差。

背景:智能体技能作为可复用的能力模块在AI系统中日益普及,医疗研究场景下的技能需超越通用评估,具备科学完整性、方法有效性、可重复性及边界安全等保障。本研究开发并初步评估了一种面向医疗研究领域的专用审计框架,重点考察其与专家评审的一致性。方法:提出MedSkillAudit([email protected])框架,分层评估技能发布前的成熟度。对五个医疗研究类别(每类15项,共75项)进行评估,两名专家独立打分(0-100)、判定发布等级(生产就绪/有限发布/测试版/拒绝)并标记高风险失败项。使用ICC(2,1)和加权卡帕系数量化系统-专家一致性,并与人类评者间一致性基线对比。结果:平均共识质量得分为72.4(标准差13.0),57.3%的技能低于有限发布阈值。系统达成ICC(2,1) = 0.449(95%置信区间:0.250-0.610),高于人类评者间一致性(ICC=0.300)。系统共识得分差异(标准差9.5)小于专家间差异(标准差12.4),且无方向性偏差(威尔科森检验p=0.613)。协议设计类技能一致性最强(ICC=0.551),学术写作类呈现负相关(ICC=-0.567),反映评分量表与专家认知不匹配。结论:领域专用的预部署审计可为医疗研究智能体技能提供实用治理基础,通过结构化审计流程弥补通用质量检查的不足。

原文摘要 · Abstract (English)

Background: Agent skills are increasingly deployed as modular, reusable capability units in AI agent systems. Medical research agent skills require safeguards beyond general-purpose evaluation, including scientific integrity, methodological validity, reproducibility, and boundary safety. This study developed and preliminarily evaluated a domain-specific audit framework for medical research agent skills, with a focus on reliability against expert review. Methods: We developed MedSkillAudit ([email protected]), a layered framework assessing skill release readiness before deployment. We evaluated 75 skills across five medical research categories (15 per category). Two experts independently assigned a quality score (0-100), an ordinal release disposition (Production Ready / Limited Release / Beta Only / Reject), and a high-risk failure flag. System-expert agreement was quantified using ICC(2,1) and linearly weighted Cohen's kappa, benchmarked against the human inter-rater baseline. Results: The mean consensus quality score was 72.4 (SD = 13.0); 57.3% of skills fell below the Limited Release threshold. MedSkillAudit achieved ICC(2,1) = 0.449 (95% CI: 0.250-0.610), exceeding the human inter-rater ICC of 0.300. System-consensus score divergence (SD = 9.5) was smaller than inter-expert divergence (SD = 12.4), with no directional bias (Wilcoxon p = 0.613). Protocol Design showed the strongest category-level agreement (ICC = 0.551); Academic Writing showed a negative ICC (-0.567), reflecting a structural rubric-expert mismatch. Conclusions: Domain-specific pre-deployment audit may provide a practical foundation for governing medical research agent skills, complementing general-purpose quality checks with structured audit workflows tailored to scientific use cases.

智能体评估医疗AI审计框架可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。