AI语言技术在多语种医疗中存在安全与公平隐患,亟需跨领域协同改进。
Artificial intelligence language technologies in multilingual healthcare: Grand challenges ahead
- 从人机协作视角审视AI医疗语言工具的性能差异与风险
- 多语言、多场景下输出可靠性不一,可能引发临床误判
- 适合研究者、政策制定者及医疗AI设计者参考
AI语言技术(AILTs)借助大语言模型(LLMs)日益融入多语种医疗工作流程,用于翻译、重写、文档生成、口译及信息传递。然而,流畅输出并不等于临床安全或公平沟通:不同语言、口音、任务和工作流中的表现差异显著,效率提升可能掩盖错误、降低可追溯性,并转移临床医生、翻译人员、口译员与医疗系统间的责任。本文综述近期同行评审研究,涵盖书面、口语及新兴代理型工作流,基于人中心AI语言技术(HCAILT)框架,分析能力、评估方法、实施模式与常见错误,聚焦可靠性、安全文化与可信度。识别文献中的关键共识与矛盾,提出未来研究与部署的七项重大挑战。我们认为,进展不仅需更优模型,还需可问责的社会技术设计、精准的人类监督,以及机器翻译/自然语言处理、翻译学、人机交互、临床实践、实施科学与政策领域的深度协作。
原文摘要 · Abstract (English)
AI language technologies (AILTs), increasingly enabled by large language models (LLMs), are becoming embedded in multilingual healthcare workflows for translation, rewriting, documentation, interpreting, and messaging in language-discordant settings. Yet fluent output is not the same as clinically safe or equitable communication: performance varies across languages, accents, tasks, and workflows, and efficiency gains can hide errors, reduce traceability, and shift responsibility across clinicians, translators, interpreters, and health systems. This narrative review synthesises recent peer-reviewed evidence across written communication, spoken communication, and emerging agentic workflows. Using the Human-Centered AI Language Technology (HCAILT) lens, it examines capabilities, evaluation practices, implementation patterns, and recurrent errors through reliability, safety culture, and trustworthiness. We identify key convergences and contradictions in the literature and propose seven grand challenges for the next phase of research and deployment. Progress, we argue, requires not only better models but also accountable sociotechnical design, calibrated human oversight, and stronger collaboration across MT/NLP, translation studies, HCI, clinical practice, implementation science, and policy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。