arXiv:2512.10780cs.CLcs.LG2025-12被引 7

罗马化输入让大模型在母婴健康分诊中出错率飙升,最高达24个百分点。

Script Gap: Evaluating LLM Triage on Indian Languages in Native vs Romanized Scripts in a Real World Setting

  • 用不确定性路由选择性处理罗马化文本,缓解脚本差异问题
  • 实测显示罗马化导致性能下降最多达24点,影响严重
  • 适合关注AI医疗安全与多语言应用的研究者和开发者

大型语言模型(LLMs)在印度高风险临床场景中日益广泛应用。印度语使用者常使用罗马化文字而非本地文字交流,但现有研究极少量化或评估这种书写方式在真实场景中的影响。我们研究了罗马化对关键领域——母婴健康分诊中LLM可靠性的影响。在涵盖五种印度语言及尼泊尔语的用户生成健康查询真实数据集上,对主流LLM进行基准测试。结果表明,罗马化文本导致性能持续下降,跨语言和模型的差距最高达24分。我们提出并评估了一种基于不确定性的选择性路由方法以缩小这一脚本差距。仅在合作的母婴健康机构,该差距可能导致近200万次额外误判。研究揭示了基于LLM的健康系统中一个关键的安全盲区:看似理解罗马化输入的模型,仍可能无法可靠响应。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly deployed in high-stakes clinical applications in India. Speakers of Indian languages frequently communicate using romanized text rather than native scripts, yet existing research rarely quantifies or evaluates this orthographic variation in real world applications. We investigate how romanization impacts the reliability of LLMs in a critical domain: maternal and newborn healthcare triage. We benchmark leading LLMs on a real world dataset of user-generated health queries spanning five Indian languages and Nepali. Our results reveal consistent degradation in performance for romanized messages, with gap reaching up to 24 points across languages and models. We propose and evaluate an Uncertainty-based Selective Routing method to close this script gap. At our partner maternal health organization alone, this gap could cause nearly 2 million excess errors in triage. Our findings highlight a critical safety blind spot in LLM-based health systems: models that appear to understand romanized input may still fail to act on it reliably.

大模型医疗AI多语言罗马化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。