评测大模型对残障群体的包容性,发现多数模型忽视特定残障群体。
Who Gets Left Behind? Auditing Disability Inclusivity in Large Language Models
- 构建覆盖九类残障的可验证评测基准,系统评估模型包容性。
- 17个模型中语音、遗传发育等残障类别支持严重不足。
- 提出兼顾广度、平衡与深度的评测方法,指导模型改进。
大型语言模型(LLMs)被越来越多地用于辅助无障碍指导,但许多残障群体仍得不到充分支持。为弥补这一差距,我们提出了一个与分类体系对齐的基准,包含经人工验证的通用无障碍问题,旨在系统性地审计不同残障类型的包容性。该基准从三个维度评估模型:问题层面覆盖率(答案内部的广度)、残障层面覆盖率(九类残障间的平衡性)和深度(支持的具体程度)。对17个专有及开源模型的应用分析显示,视觉、听觉和运动功能障碍常被覆盖,而语音、遗传/发育、感官认知和心理健康相关需求则显著被忽视。深度支持也集中在少数类别,其余类别普遍薄弱。研究揭示了当前大模型在无障碍指导中的遗漏群体,并提出可操作的改进路径:采用分类意识的提示设计或训练,以及联合评估广度、平衡与深度的评测机制。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly used for accessibility guidance, yet many disability groups remain underserved by their advice. To address this gap, we present taxonomy aligned benchmark1 of human validated, general purpose accessibility questions, designed to systematically audit inclusivity across disabilities. Our benchmark evaluates models along three dimensions: Question-Level Coverage (breadth within answers), Disability-Level Coverage (balance across nine disability categories), and Depth (specificity of support). Applying this framework to 17 proprietary and open-weight models reveals persistent inclusivity gaps: Vision, Hearing, and Mobility are frequently addressed, while Speech, Genetic/Developmental, Sensory-Cognitive, and Mental Health remain under served. Depth is similarly concentrated in a few categories but sparse elsewhere. These findings reveal who gets left behind in current LLM accessibility guidance and highlight actionable levers: taxonomy-aware prompting/training and evaluations that jointly audit breadth, balance, and depth.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。