盲人用户通过自定义提示词改善对话式视觉问答体验。
Say It My Way: Exploring Control in Conversational Visual Question Answering with Blind Users
- 盲人用户通过提示工程绕过系统限制,实现个性化交互。
- 平均每次对话3轮,最长达21轮,输入文本仅为回应的十分之一。
- 研究为无障碍交互设计提供实证依据,适合残障辅助技术开发者。
提示与引导技术在通用生成式AI中已成熟,但面向视障用户的辅助型视觉问答(VQA)工具仍采用固定交互模式,定制化能力有限。当系统响应与用户目标和上下文不一致时,用户控制尤为重要,这一差距对依赖此类系统获取信息的视障人群尤为关键。我们邀请11位视障用户参与真实世界对话式VQA系统的交互定制。基于418次交互、反思记录及访谈,分析了参与者采用的提示技术,包括研究中引入的方法及他们在实际使用中独立开发的技巧。研究发现,对话通常较长:平均3轮,最多21轮;输入文本长度约为系统回复的十分之一。尽管系统基于最先进的大语言模型(LLM),但仍缺乏语气控制、空间/时间距离估计能力弱、依赖不可访问的图像构图,且几乎没有相机引导。我们探讨了提示工程等定制化技术如何帮助用户克服这些局限。除提出一个公开可用的新数据集外,研究还为查询层面与系统层面的交互设计提供了重要启示。
原文摘要 · Abstract (English)
Prompting and steering techniques are well established in general-purpose generative AI, yet assistive visual question answering (VQA) tools for blind users still follow rigid interaction patterns with limited opportunities for customization. User control can be helpful when system responses are misaligned with their goals and contexts, a gap that becomes especially consequential for blind users that may rely on these systems for access. We invite 11 blind users to customize their interactions with a real-world conversational VQA system. Drawing on 418 interactions, reflections, and post-study interviews, we analyze prompting-based techniques participants adopted, including those introduced in the study and those developed independently in real-world settings. VQA interactions were often lengthy: participants averaged 3 turns, sometimes up to 21, with input text typically tenfold shorter than the responses they heard. Built on state-of-the-art LLMs, the system lacked verbosity controls, was limited in estimating distance in space and time, relied on inaccessible image framing, and offered little to no camera guidance. We discuss how customization techniques such as prompt engineering can help participants work around these limitations. Alongside a new publicly available dataset, we offer insights for interaction design at both query and system levels.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。