首个面向视障辅助的VLM评判基准,揭示现有模型评估不可靠。
VIABLE: A Visually Impaired Assistance Benchmark for VLM-as-a-Judge Evaluation

- 构建视障辅助专用评判基准VIABLE,覆盖3大场景30万+样本。
- 7个模型评测显示最强仅52.6%故障诊断准确率,且存在严重偏见。
- 提出VIA-Judge-Agent增强框架,提升诊断与用户偏好响应。
基于AI的视障辅助(VIA)仍面临人工评估成本高的挑战。尽管视觉语言模型作为评判者(VLM-as-a-Judge)在通用领域展现出潜力,但在VIA任务中是否可信尚不明确。为此,我们提出首个针对VIA任务的VLM评判基准——VIABLE(Visually Impaired Assistance Benchmark for VLM-as-a-Judge Evaluation)。VIABLE包含跨三个场景的超30万条判断样本,并引入有效性-公正性-稳定性评估框架及12类故障分类体系。基于此,我们对七个不同规模模型的系统性研究发现:现有模型在所有评估维度上均不可靠。最强模型GPT-5.4的单故障诊断准确率仅为52.6%,但自偏好率达94.2%;开源模型则表现出强烈偏见和对抗脆弱性。为解决上述问题,我们提出VIA-Judge-Agent——一种无需修改模型的推理时增强框架,通过视觉证据提取与分类引导工作流,显著提升诊断准确率及盲人用户更偏好的下游辅助响应。数据与代码已开源:https://github.com/YiyiyiZhao/VIABLE。
原文摘要 · Abstract (English)
AI-based Visually Impaired Assistance (VIA) remains challenging, largely due to the high cost of human evaluation. The VLM-as-a-Judge paradigm may offer a promising alternative, although it has mostly been studied in general domains. We therefore ask whether such judges can be trusted for VIA tasks. To investigate this question, we introduce VIABLE (Visually Impaired Assistance Benchmark for VLM-as-a-Judge Evaluation), the first benchmark for VLM-as-a-Judge evaluation in VIA. VIABLE contains over 300K judgment samples across three scenarios and introduces an Effectiveness--Impartiality--Stability framework with a 12-mode failure taxonomy. Based on VIABLE, our systematic study of seven judges across different model scales shows that existing models are largely unreliable across all evaluation axes. The strongest judge, GPT-5.4, achieves only 52.6% single-failure diagnostic accuracy, yet exhibits the highest self-preference rate at 94.2%; while open-source judges are strongly biased and adversarially fragile. To address these issues, we propose VIA-Judge-Agent, a model-agnostic inference-time harness that augments judges with visual evidence extraction and a taxonomy-guided workflow. It enables positive improvements in diagnostic accuracy and downstream VIA responses more preferred by BLV users. Data and code are available at: https://github.com/YiyiyiZhao/VIABLE
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。