测试多模态大模型在灾情救助中跨文本与音频的响应一致性
Evaluating Multimodal LLMs across Text and Audio Modalities for Accessible Disaster Assistance

- 用真实灾情场景测试4类弱势人群的跨模态输出一致性
- 所有模型均无法保证文本与音频响应一致,弱势群体差距更大
- 为无障碍灾情沟通系统设计提供公平性改进方向
有效的灾害风险传播是人道主义的核心挑战,但现有应急体系未能满足听力障碍者、孕妇、带幼儿母亲及失智老年人等有特殊需求人群的需求。近年来,多模态大语言模型(MM-LLMs)在单一系统中整合文本、音频、图像和视频处理能力,展现出服务多元用户(如聊天机器人)的强大潜力。然而,其部署适用性仍依赖一个被忽视的关键属性:用户通过不同模态(如文字或语音)输入时,系统是否能生成一致且可操作的输出。本文针对四种典型脆弱人群,在真实灾情预警场景下,全面评估开源多模态大模型在文本与音频模态间的响应一致性。结果表明,无一模型能实现跨模态可靠一致,且对有特殊需求的人群性能差距显著扩大,导致模态依赖型不公平,削弱了这些系统的人道价值。研究据此提出具体设计建议,推动构建更公平、可信、包容的灾害风险沟通AI工具。
原文摘要 · Abstract (English)
Effective disaster risk communication is a foundational humanitarian challenge, yet current emergency infrastructure fails to meet the needs of individuals with access and functional needs, including hard-of-hearing individuals, pregnant women, mothers with toddlers, and elderly individuals with dementia. Recent advancements in Artificial Intelligence (AI), especially Multi-Modal Large Language Models (MM-LLMs), demonstrate powerful capabilities to serve diverse users across text, audio, image, and video modalities within a single unified system, such as a chatbot. However, their suitability for deployment rests on a property that receives limited scrutiny, i.e., whether these systems produce consistent, actionable outputs regardless of the modality through which a user communicates. In this paper, we conduct a comprehensive analysis to understand the status of open-weight MM-LLMs using real emergency alert scenarios across four different vulnerable personas. These state-of-the-art (SOTA) models are evaluated on consistency of responses across text and audio modalities when the same task scenario is given. Findings indicate that no model achieves reliable consistency across modalities, and that performance gaps are heightened for personas with access needs, introducing modality-dependent inequity that undermines the humanitarian value of these systems. These results inform concrete design recommendations for building equitable, trustworthy, and inclusive AI tools for disaster risk communication.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。