arXiv:2508.18328cs.CLcs.CY2025-08被引 4

发现网页多语言内容与辅助技术不匹配,导致视障用户访问困难

Not All Visitors are Bilingual: A Measurement Study of the Multilingual Web from an Accessibility Perspective

  • 构建120,000个非拉丁语系网站的大规模数据集LangCrUX
  • 87%的网页语言提示与可见内容语言不一致,影响屏幕阅读器识别
  • 提出Kizuki工具,自动检测并修复语言不一致的可访问性问题

英语主导网络,占全球前十万网站近一半。尽管越来越多网站混合使用英语与本地语言(包括隐藏元数据),但辅助技术如屏幕阅读器对非拉丁文字支持不足,常导致误读或误发音,加剧视障用户在多语言环境中的访问障碍。现有研究受限于缺乏大规模多语言网页数据。为此,我们构建了首个涵盖12种主要非拉丁语言的12万网站数据集LangCrUX。基于此,系统分析发现:多数网页的语言提示未反映可见内容的语言多样性,严重削弱屏幕阅读器效果。最后提出Kizuki——一种能感知语言的自动化可访问性测试扩展,以应对语言不一致提示带来的局限性。

原文摘要 · Abstract (English)

English is the predominant language on the web, powering nearly half of the world's top ten million websites. Support for multilingual content is nevertheless growing, with many websites increasingly combining English with regional or native languages in both visible content and hidden metadata. This multilingualism introduces significant barriers for users with visual impairments, as assistive technologies like screen readers frequently lack robust support for non-Latin scripts and misrender or mispronounce non-English text, compounding accessibility challenges across diverse linguistic contexts. Yet, large-scale studies of this issue have been limited by the lack of comprehensive datasets on multilingual web content. To address this gap, we introduce LangCrUX, the first large-scale dataset of 120,000 popular websites across 12 languages that primarily use non-Latin scripts. Leveraging this dataset, we conduct a systematic analysis of multilingual web accessibility and uncover widespread neglect of accessibility hints. We find that these hints often fail to reflect the language diversity of visible content, reducing the effectiveness of screen readers and limiting web accessibility. We finally propose Kizuki, a language-aware automated accessibility testing extension to account for the limited utility of language-inconsistent accessibility hints.

网页无障碍多语言辅助技术数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。