arXiv:2511.13458cs.HCcs.AI2025-11

通过用户工作坊研究视觉语言模型的信任形成机制。

Trust in Vision-Language Models: Insights from a Participatory User Workshop

  • 采用用户中心方法开展参与式工作坊,探索信任构建过程。
  • 发现用户对VLM的信任受上下文和交互体验影响显著。
  • 为未来设计可信评估体系提供实证基础,适合人机交互研究者。

随着视觉语言模型(VLMs)在大规模图像-文本和视频-文本数据集上预训练并日益部署,亟需为用户提供判断何时信赖这些系统的能力。然而,用户对VLMs的信任如何建立与演变仍是未解问题。这一挑战因越来越多地依赖AI模型作为实验验证的评判者而加剧,以规避直接开展用户参与式设计研究的成本与复杂性。本文采用用户中心方法,报告了一次面向潜在VLM用户的工作坊的初步结果。该试点研究的洞见将指导未来研究,以实现信任度量的上下文适配,并制定参与者参与策略,更好地契合用户与VLM的交互场景。

原文摘要 · Abstract (English)

With the growing deployment of Vision-Language Models (VLMs), pre-trained on large image-text and video-text datasets, it is critical to equip users with the tools to discern when to trust these systems. However, examining how user trust in VLMs builds and evolves remains an open problem. This problem is exacerbated by the increasing reliance on AI models as judges for experimental validation, to bypass the cost and implications of running participatory design studies directly with users. Following a user-centred approach, this paper presents preliminary results from a workshop with prospective VLM users. Insights from this pilot workshop inform future studies aimed at contextualising trust metrics and strategies for participants' engagement to fit the case of user-VLM interaction.

视觉语言模型用户信任人机交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。