arXiv:2504.07423cs.HCcs.AI2025-04中稿 · the CHI '25 Worksh…被引 3

提醒医疗AI评估别只看信任和依赖,要真实反映临床价值。

Over-Relying on Reliance: Towards Realistic Evaluations of AI-Based Clinical Decision Support

  • 反对仅用信任、依赖等指标评价医疗AI,主张更真实的评估方式。
  • 指出当前评估忽视了AI在实际临床中的局限与医生的灵活使用。
  • 适合关注医疗AI设计与真实场景评估的研究者和开发者。

随着基于人工智能的临床决策支持(AI-CDS)在医疗服务中日益普及,人机交互(HCI)研究在设计人工智能与临床医生协同工作方面扮演着越来越重要的角色。然而,目前对AI-CDS的评估往往未能准确捕捉到人工智能在哪些情境下对临床医生有用,哪些情境下无用。本文反思了我们的工作及具有影响力的AI-CDS文献,呼吁跳出以信任、依赖、接受度和任务性能为核心的评估框架(我们称之为‘人机协作陷阱’)。尽管这些指标在某些简单场景中具有一定意义,但我们认为,过度追求这些指标会忽略人工智能在临床实践中未能实现的价值,以及医生成功利用人工智能的方式。随着人机交互与医疗人工智能领域不断发展新的设计与评估方法,我们呼吁学术界优先采用生态有效、符合医疗领域的研究范式,以衡量人工智能为医护人员带来的涌现性价值。

原文摘要 · Abstract (English)

As AI-based clinical decision support (AI-CDS) is introduced in more and more aspects of healthcare services, HCI research plays an increasingly important role in designing for complementarity between AI and clinicians. However, current evaluations of AI-CDS often fail to capture when AI is and is not useful to clinicians. This position paper reflects on our work and influential AI-CDS literature to advocate for moving beyond evaluation metrics like Trust, Reliance, Acceptance, and Performance on the AI's task (what we term the "trap" of human-AI collaboration). Although these metrics can be meaningful in some simple scenarios, we argue that optimizing for them ignores important ways that AI falls short of clinical benefit, as well as ways that clinicians successfully use AI. As the fields of HCI and AI in healthcare develop new ways to design and evaluate CDS tools, we call on the community to prioritize ecologically valid, domain-appropriate study setups that measure the emergent forms of value that AI can bring to healthcare professionals.

医疗AI人机协作评估方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。