arXiv:2502.15226cs.CLcs.AI2025-02ACL被引 6

用AI实时采访用户,挖掘对大模型的真实看法。

Understand User Opinions of Large Language Models via LLM-Powered In-the-Moment User Experience Interviews

  • 用LLM充当访谈员,交互后即时收集用户反馈。
  • 发现用户对DeepSeek-R1推理过程评价两极分化。
  • 适合研究人机交互、用户体验的学者与开发者。

大模型哪个更好?每项评测都有其故事,但用户真实想法是什么?本文提出CLUE——一个基于大模型的访谈系统,能在用户与大模型交互后立即开展实时体验访谈,并自动从海量访谈记录中提炼用户观点。我们招募数千名用户,先与目标大模型对话,再由CLUE进行访谈。实验表明,CLUE捕捉到有趣见解,例如用户对DeepSeek-R1显示推理过程存在明显两极化评价,同时普遍要求信息时效性和多模态支持。代码与数据已公开于https://github.com/cxcscmu/LLM-Interviewer。

原文摘要 · Abstract (English)

Which large language model (LLM) is better? Every evaluation tells a story, but what do users really think about current LLMs? This paper presents CLUE, an LLM-powered interviewer that conducts in-the-moment user experience interviews, right after users interact with LLMs, and automatically gathers insights about user opinions from massive interview logs. We conduct a study with thousands of users to understand user opinions on mainstream LLMs, recruiting users to first chat with a target LLM and then be interviewed by CLUE. Our experiments demonstrate that CLUE captures interesting user opinions, e.g., the bipolar views on the displayed reasoning process of DeepSeek-R1 and demands for information freshness and multi-modality. Our code and data are at https://github.com/cxcscmu/LLM-Interviewer.

用户研究大模型访谈系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。