arXiv:2503.15512cs.HCcs.AI2025-03

开发者缺乏培训,仅靠政策无法做出对用户有意义的AI解释。

Policy alone is probably not the solution: A large-scale experiment on how developers struggle to design meaningful end-user explanations

  • 194名开发者在无训练情况下设计医疗AI解释,仅靠政策指导效果差。
  • 政策执行加强后合规率提高,但解释仍不适合医生和患者使用。
  • 多数解释技术错误、术语过多,且以开发者为中心,缺乏临床场景支持。

开发者在实际中决定机器学习系统的解释方式,但他们通常未接受面向非技术用户的解释设计培训。尽管透明度和可解释性要求日益被法规和组织政策规定,但这些政策如何影响开发者行为或解释质量仍不明确。我们报告了两项针对194名典型开发者的受控实验,他们为一个用于糖尿病视网膜病变筛查的ML工具设计解释,均无专门的人本可解释AI训练。第一项实验显示,政策目的与细节差异影响甚微:政策常被忽略,解释质量始终较低。第二项实验中,更强的执行提高了形式合规性,但解释仍大多不适合医疗专业人员和患者。进一步观察发现,开发者反复产出技术缺陷、难以理解、面向开发者而非终端用户、依赖医学术语、未充分结合临床决策背景与流程的解释,其中开发者中心化表述最为普遍。结果表明,仅靠政策与执行不足以生成有意义的终端用户解释,负责任AI框架可能高估了开发者在无额外培训、工具或实施支持下,将高层要求转化为以人为本设计的能力。

原文摘要 · Abstract (English)

Developers play a central role in determining how machine learning systems are explained in practice, yet they are rarely trained to design explanations for non-technical audiences. Despite this, transparency and explainability requirements are increasingly codified in regulation and organizational policy. It remains unclear how such policies influence developer behavior or the quality of the explanations they produce. We report results from two controlled experiments with 194 participants, typical developers without specialized training in human-centered explainable AI, who designed explanations for an ML-powered diabetic retinopathy screening tool. In the first experiment, differences in policy purpose and level of detail had little effect: policy guidance was often ignored and explanation quality remained low. In the second experiment, stronger enforcement increased formal compliance, but explanations largely remained poorly suited to medical professionals and patients. We further observed that across both experiments, developers repeatedly produced explanations that were technically flawed or difficult to interpret, framed for developers rather than end users, reliant on medical jargon, or insufficiently grounded in the clinical decision context and workflow, with developer-centric framing being the most prevalent. These findings suggest that policy and policy enforcement alone are insufficient to produce meaningful end-user explanations and that responsible AI frameworks may overestimate developers' ability to translate high-level requirements into human-centered designs without additional training, tools, or implementation support.

AI解释开发者医疗AI人因设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。