arXiv:2607.00019cs.CYcs.AI2026-07ACL

测试大模型在紧急呼叫系统中的实际表现,揭示技术应用的潜在风险。

LLMs in the Real World: Evaluating "AI" in Emergency Contexts

  • 构建支持55种语言的文本转911系统,评估真实场景下的AI可靠性
  • 发现当前部署中存在对技术能力的普遍误解,导致应急响应风险
  • 提出从开发到部署各环节的实用建议,适合政策制定者与技术团队参考

本文呼吁研究界更积极地向公众传达研究成果。为说明问题严重性,我们以一个基于大模型的机器翻译应用在真实场景中的初步部署为例:一个支持55种语言的文本转911系统,旨在紧急情况下帮助无法直接拨打电话的用户。我们识别出对这类技术的若干常见误解,并据此提出一系列针对开发与部署全链条利益相关方的具体建议与最佳实践。尽管科研进展常聚焦于解决‘难题’,但我们认为,那些看似‘简单’——即现有技术已足够应对——的问题反而最易被忽视。

原文摘要 · Abstract (English)

This paper offers a call to action. We urge our colleagues in the research community to play a greater role in the articulation of our findings to the public. To illustrate the stakes we present a case study on the initial stages of an LLM-based machine translation application's deployment in a real-world context: a text-2-911 system advertising capabilities in 55 languages for use in emergencies in which it may be difficult to call operators directly. We identify a number of common misconceptions about technologies such as these, concluding with a set of concrete recommendations and best practices for stakeholders at every stage of the development and deployment pipeline. While the advancement of scientific research often lies in solving the "hard" problems, we argue it is often the "easy" ones -- problems for which the latest technology is often unnecessary -- that are most overlooked.

大模型应用紧急响应人机交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。