arXiv:2505.05318cs.CVcs.AI2025-05被引 3

梳理视觉语言模型用户信任机制,指明研究方向与挑战。

Mapping User Trust in Vision Language Models: Research Landscape, Challenges, and Prospects

  • 构建跨学科信任分析框架,涵盖认知能力与交互模式。
  • 基于用户调研提出未来研究的初步需求清单。
  • 适合关注AI可信性与人机协作的研究者参考。

视觉语言模型(VLMs)在大规模图像-文本和视频-文本数据集上预训练后迅速普及,亟需保护并告知用户何时应信任这些系统。本文综述了用户与VLM交互中信任动态的相关研究,通过多学科分类体系涵盖不同认知科学能力、协作模式及代理行为。结合研讨会中潜在用户的意见,提炼出未来VLM信任研究的初步要求。

原文摘要 · Abstract (English)

The rapid adoption of Vision Language Models (VLMs), pre-trained on large image-text and video-text datasets, calls for protecting and informing users about when to trust these systems. This survey reviews studies on trust dynamics in user-VLM interactions, through a multi-disciplinary taxonomy encompassing different cognitive science capabilities, collaboration modes, and agent behaviours. Literature insights and findings from a workshop with prospective VLM users inform preliminary requirements for future VLM trust studies.

视觉语言模型用户信任人机交互综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。