arXiv:2510.17764cs.CL2025-10综述被引 1

用自主等级框架重新评估医疗大模型,推动临床落地

Evaluating Medical LLMs by Levels of Autonomy: A Survey Moving from Benchmarks to Applications

  • 按自主程度分4级(L0-L3)重构评估体系
  • 明确每级对应的任务与风险,对齐评估指标
  • 适合关注临床安全与可信部署的研究者

医疗大语言模型在标准基准测试中表现优异,但其在临床工作流中的安全可靠应用仍面临挑战。本综述提出以自主等级(L0-L3)为评估视角,涵盖信息工具、信息处理与整合、决策支持及受控代理四类角色。将现有基准与指标与各等级允许的操作及其风险相匹配,使评估目标更加明确。据此提出分级的评估蓝图,指导指标选择、证据构建与结论报告,并连接评估与监管机制。通过聚焦自主性,推动领域从单纯分数宣称转向支持真实临床应用的可信、风险敏感型证据。

原文摘要 · Abstract (English)

Medical Large language models achieve strong scores on standard benchmarks; however, the transfer of those results to safe and reliable performance in clinical workflows remains a challenge. This survey reframes evaluation through a levels-of-autonomy lens (L0-L3), spanning informational tools, information transformation and aggregation, decision support, and supervised agents. We align existing benchmarks and metrics with the actions permitted at each level and their associated risks, making the evaluation targets explicit. This motivates a level-conditioned blueprint for selecting metrics, assembling evidence, and reporting claims, alongside directions that link evaluation to oversight. By centering autonomy, the survey moves the field beyond score-based claims toward credible, risk-aware evidence for real clinical use.

医疗AI评估框架自主等级

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。