给AI装上医学分析技能包,能提升肺癌基因组研究的输出质量。
Skill-Augmented AI Agents for Medical Research Analysis: An Exploratory Multi-Model Human Evaluation in an NSCLC Transcriptomic Biomarker Task

- 让AI自主调用医学研究技能包完成分析任务
- 技能增强版输出获专家评分更高(均值5.50 vs 5.11)
- 适合关注AI在生物医学研究中可靠性的研究者
大型语言模型和AI代理在生物医学研究中应用日益广泛,但其原生输出可能遗漏关键分析步骤、误用方法或过度推论。我们评估了自主访问医学研究技能包是否能提升AI生成的非小细胞肺癌免疫治疗生物标志物分析结果的质量。采用非小细胞肺癌免疫治疗生物标志物任务,测试六种模型骨干。共生成21份匿名输出:9份原生AI输出与12份通过OpenClaw实现的技能增强输出。四位非专家生物医学评审员与两位盲评专家各对每份输出进行两次评分。主要结局为专家评定的整体质量。结果显示,技能增强输出的专家整体质量方向性更高(均值5.50对比5.11;差值=0.39;自举95%置信区间-0.04至0.90;Welch p=0.156)。非专家评分也呈现相同趋势(均值4.72对比4.47;差值=0.26;自举95%置信区间-0.25至0.80;Welch p=0.373)。专家一致性较低(单次评分组内相关系数ICC=-0.15),模型特定效应描述性且异质。结论:在本探索性样本中,自主技能访问显示出质量提升的倾向信号,但该信号小于专家评分噪声,不应视为确认证据。研究主要推动更大规模、具更强可靠性控制、平台可复现性及生物学有效性评估的技能增强型AI代理研究。
原文摘要 · Abstract (English)
Background. Large language models and AI agents are increasingly used to support biomedical research, but native model outputs may omit key analytical steps, misuse methods, or overstate conclusions. We evaluated whether autonomous access to a medical research skill package was associated with higher-quality AI-generated transcriptomic research-analysis outputs compared with native AI without skills. Methods. We conducted an exploratory multi-model human evaluation using a non-small cell lung cancer immunotherapy biomarker task. Six model backbones were tested. The evaluation included 21 anonymized outputs: 9 native-AI outputs and 12 skill-augmented outputs generated through an AI agent implementation represented by OpenClaw. Four non-expert biomedical reviewers and two blinded experts evaluated each output, with two ratings from each reviewer type. The primary outcome was expert-rated overall quality. Results. Skill-augmented outputs showed directionally higher expert overall quality than native-AI outputs (mean 5.50 vs 5.11; difference=0.39; bootstrap 95\% CI, -0.04 to 0.90; Welch p=0.156). Non-expert reviewer quality showed the same direction (mean 4.72 vs 4.47; difference=0.26; bootstrap 95\% CI, -0.25 to 0.80; Welch p=0.373). Expert agreement was limited (single-rating ICC=-0.15), and model-specific effects were descriptive and heterogeneous. Conclusions. Autonomous skill access showed a directional quality signal in this exploratory sample, but the signal was smaller than expert-rating noise and should not be interpreted as confirmatory evidence. The findings primarily motivate larger evaluations of skill-augmented AI agents with stronger reliability controls, platform replication, and biological-validity assessment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。