用大模型检测手机应用屏幕阅读器无障碍缺陷,覆盖率达69.2%。
ScreenAudit: Detecting Screen Reader Accessibility Errors in Mobile Apps Using Large Language Models
- 用大模型遍历界面,分析元数据与文本,识别遗漏错误
- 相比主流检查工具,错误覆盖率从31.3%提升至69.2%
- 专家认可其反馈质量高,适合开发者实际使用
许多移动应用存在无障碍缺陷,使残障用户无法使用。现有基于规则的检查工具虽能早期发现部分问题,但检测类型有限。我们提出ScreenAudit,一个基于大语言模型的系统,通过遍历移动应用界面,提取元数据和文本内容,识别屏幕阅读器相关的未被现有工具覆盖的无障碍错误。我们邀请六位无障碍专家(含一名屏幕阅读器使用者)对14个不同应用界面的报告进行评估。结果显示,ScreenAudit平均错误覆盖率达69.2%,远超主流检查工具仅31.3%的覆盖率。专家反馈表明,ScreenAudit提供的反馈质量更高,涵盖更多屏幕阅读器无障碍维度,且在真实开发场景中具有实用价值。
原文摘要 · Abstract (English)
Many mobile apps are inaccessible, thereby excluding people from their potential benefits. Existing rule-based accessibility checkers aim to mitigate these failures by identifying errors early during development but are constrained in the types of errors they can detect. We present ScreenAudit, an LLM-powered system designed to traverse mobile app screens, extract metadata and transcripts, and identify screen reader accessibility errors overlooked by existing checkers. We recruited six accessibility experts including one screen reader user to evaluate ScreenAudit's reports across 14 unique app screens. Our findings indicate that ScreenAudit achieves an average coverage of 69.2%, compared to only 31.3% with a widely-used accessibility checker. Expert feedback indicated that ScreenAudit delivered higher-quality feedback and addressed more aspects of screen reader accessibility compared to existing checkers, and that ScreenAudit would benefit app developers in real-world settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。