首个评估多模态AI辅助无障碍出行规划的基准测试
MAP: A Benchmark on Multimodal Accessibility Planning for Real World Places

- 构建双评估体系:验证无障碍声明与检索视觉证据
- 支持动态更新真实世界场所信息,确保评测时效性
- 适合研究无障碍智能助手与多模态推理的学者
我们提出MAP,首个用于评估多模态AI系统在现实世界场所访问规划中为有无障碍需求用户提供建议能力的基准。在评估中,系统需验证或推荐满足特定无障碍要求的景点。MAP包含两项新评估:无障碍规划中的声明验证,检测场所信息与声明的无障碍特性是否一致,并识别满足要求的场所;以及无障碍规划中的视觉证据检索,检验多模态AI能否为指定场所和无障碍特征选择合适的视觉证据。该方法通过定期刷新真实数据与标注,支持对系统在随时间变化的信息环境下的持续评估。基准采用自动评分与人工评分相结合的方式,对部分结果进行人工校验。
原文摘要 · Abstract (English)
We introduce MAP, the first benchmark to evaluate multimodal AI systems as assistants for users with accessibility requirements when planning visits to places in the real world. In our evaluation, systems are presented with requests to verify or recommend a point of interest meeting an accessibility requirement. MAP contains two novel assessments: Claim verification for accessibility planning assesses if information on places and stated accessibility features is supported and identifies places that satisfy requested accessibility features. Visual evidence retrieval for accessibility planning checks if a multimodal AI system can select visual evidence for the requested place and accessibility feature. Our methodology supports comparison of AI systems in a setting where place information and accessibility information can change over time by evaluating systems and refreshing ground truth data at scheduled times. The benchmark is based on automatic rating and human rating for a proportion of responses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。