用AI助手让网页可访问性审计更高效可扩展
Towards Scalable Web Accessibility Audit with MLLMs as Copilots
- 构建人机协作框架,用图结构采样确保页面覆盖全面
- 多模态大模型辅助审计,显著降低人工耗时
- 适合需要大规模网页合规检查的团队或机构
保障网页可访问性对推动数字空间的社会公平至关重要,但当前多数网站界面仍不合规,主要因审计流程资源消耗大、难以扩展。尽管WCAG-EM提供了结构化评估方法,但依赖大量人力且缺乏规模化执行支持。本文提出审计框架AAA,通过人-人工智能协同模式实现WCAG-EM的落地。核心创新包括:GRASP——基于图的多模态采样方法,利用视觉、文本与关系线索的嵌入表示实现代表性页面覆盖;MaC——基于多模态大语言模型的智能助手,支持跨模态推理和高难度任务的自动化辅助。二者结合实现了可扩展的端到端审计流程,赋能人工审计员提升真实世界影响力。此外,我们构建了四个新数据集,用于评估审计流程各阶段。大量实验表明,微调后的小规模语言模型亦可胜任专业审计任务。
原文摘要 · Abstract (English)
Ensuring web accessibility is crucial for advancing social welfare, justice, and equality in digital spaces, yet the vast majority of website user interfaces remain non-compliant, due in part to the resource-intensive and unscalable nature of current auditing practices. While WCAG-EM offers a structured methodology for site-wise conformance evaluation, it involves great human efforts and lacks practical support for execution at scale. In this work, we present an auditing framework, AAA, which operationalizes WCAG-EM through a human-AI partnership model. AAA is anchored by two key innovations: GRASP, a graph-based multimodal sampling method that ensures representative page coverage via learned embeddings of visual, textual, and relational cues; and MaC, a multimodal large language model-based copilot that supports auditors through cross-modal reasoning and intelligent assistance in high-effort tasks. Together, these components enable scalable, end-to-end web accessibility auditing, empowering human auditors with AI-enhanced assistance for real-world impact. We further contribute four novel datasets designed for benchmarking core stages of the audit pipeline. Extensive experiments demonstrate the effectiveness of our methods, providing insights that small-scale language models can serve as capable experts when fine-tuned.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。