揭示编程助手用户对大模型能力的常见误解,避免盲目依赖。
User Misconceptions of LLM-Based Conversational Programming Assistants
- 通过头脑风暴与对话分析,识别用户对工具功能的错误预期。
- 发现用户普遍高估网页搜索、代码执行和非文本输出能力。
- 适合初学者和开发者工具设计者参考,提升人机协作效率。
基于大语言模型(LLMs)的编程助手日益普及,尤其是像ChatGPT这样的对话式助手,对新手程序员尤为友好。然而,不同工具功能差异大,且扩展功能(如网络搜索、代码执行、检索增强生成)可用性不一,容易引发用户误解,导致过度依赖、低效实践或质量控制不足。本文采用两阶段方法:首先梳理潜在误解,再对来自WildChat数据集的Python编程对话进行定性分析。结果表明,用户对网页访问、代码执行及非文本输出等功能存在明显误判。此外,还暴露出调试、验证与优化过程中对信息需求的认知偏差。研究强调,需让LLM工具更清晰地传达自身能力,并在编程场景中以实证方式澄清关键概念。
原文摘要 · Abstract (English)
Programming assistants powered by large language models (LLMs) have become widely available, with conversational assistants like ChatGPT particularly accessible to novice programmers. However, varied tool capabilities and inconsistent availability of extensions (web search, code execution, retrieval-augmented generation) create opportunities for user misconceptions that may lead to over-reliance, unproductive practices, or insufficient quality control. We characterize misconceptions that users of conversational LLM-based assistants may have in programming contexts through a two-phase approach: first brainstorming and cataloging potential misconceptions, then conducting qualitative analysis of Python-programming conversations from the WildChat dataset. We find evidence that users have misplaced expectations about features like web access, code execution, and non-text outputs. We also note the potential for deeper conceptual issues around information requirements for debugging, validation, and optimization. Our findings reinforce the need for LLM-based tools to more clearly communicate their capabilities to users and empirically ground aspects that require clarification in programming contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。