arXiv:2511.08917cs.HCcs.CV2025-11被引 3

测试图像质量对视障者用AI生成商品描述的影响

"It's trained by non-disabled people": Evaluating How Image Quality Affects Product Captioning with Vision-Language Models

  • 基于视障者拍摄的1859张产品图,系统评估图像模糊等质量问题
  • 图像有缺陷时,最优模型准确率从98%降至75%,问题叠加更严重
  • 呼吁以视障者真实体验为中心改进AI模型评估与设计

视觉语言模型(VLMs)正被越来越多的盲人和低视力(BLV)人群用于识别日常生活中的商品,如食品、个人护理用品和家居用品。尽管应用广泛,我们仍缺乏对常见图像质量问题(如模糊、构图错误、旋转)如何影响VLM生成描述准确性,以及这些描述是否满足BLV人群信息需求的实证理解。基于对86名BLV参与者的调查,我们构建了一个包含1,859张来自BLV人群的真实产品图像的标注数据集,系统评估图像质量问题对VLM生成描述的影响。在无质量问题的图像上,最优VLM准确率达98%;但当存在图像质量问题时,整体准确率下降至75%,且问题越叠加,性能越差。本文强调需在整个研发过程中以残障人士体验为中心开展模型评估,并为HCI与ML研究者提供具体建议,以提升VLM对BLV人群的可靠性。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) are increasingly used by blind and low-vision (BLV) people to identify and understand products in their everyday lives, such as food, personal care items, and household goods. Despite their prevalence, we lack an empirical understanding of how common image quality issues--such as blur, misframing, and rotation--affect the accuracy of VLM-generated captions and whether the resulting captions meet BLV people's information needs. Based on a survey of 86 BLV participants, we develop an annotated dataset of 1,859 product images from BLV people to systematically evaluate how image quality issues affect VLM-generated captions. While the best VLM achieves 98% accuracy on images with no quality issues, accuracy drops to 75% overall when quality issues are present, worsening considerably as issues compound. We discuss the need for model evaluations that center on disabled people's experiences throughout the process and offer concrete recommendations for HCI and ML researchers to make VLMs more reliable for BLV people.

视觉语言模型视障辅助图像质量用户体验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。