arXiv:2604.16304cs.SEcs.AI2026-04被引 1

揭示LLM产品评估中‘结果与行动脱节’的痛点,提出实用改进策略。

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild

  • 通过访谈发现从业者从直觉判断到组织协调的十种评估实践。
  • 识别出'结果行动力缺口':数据收集但难转化为具体优化。
  • 建议系统化方法,帮助团队从随意判断转向结构化评估。

随着大型语言模型(LLMs)被集成到数字产品中,其不可预测性使得传统评估方式失效。通过对来自不同领域的19位从业者的访谈,我们识别出涵盖非正式‘直觉判断’到组织层面元工作在内的十种评估实践。除了确认已知的四个挑战外,本文提出一个新挑战——‘结果-行动力缺口’,即从业者虽能收集评估数据,却难以将其转化为具体改进措施。基于成功团队的经验,我们提出若干策略以弥合该缺口,支持从业者从非正式的解读性实践(如直觉判断)向系统化评估过渡。分析表明,这些解读性实践是应对LLM特性的必要适应,而非方法论失败。对人机交互研究者而言,这为支持从业者系统化新兴实践提供了研究机会,而非开发新评估框架。

原文摘要 · Abstract (English)

How do product teams evaluate LLM-powered products? As organizations integrate large language models (LLMs) into digital products, their unpredictable nature makes traditional evaluation approaches inadequate, yet little is known about how practitioners navigate this challenge. Through interviews with nineteen practitioners across diverse sectors, we identify ten evaluation practices spanning informal 'vibe checks' to organizational meta-work. Beyond confirming four documented challenges, we introduce a novel fifth we call the results-actionability gap, in which practitioners gather evaluation data but cannot translate findings into concrete improvements. Drawing on patterns from successful teams, we contribute strategies to bridge this gap, supporting practitioners' formalization journey from ad-hoc interpretive practices (e.g., vibe checks) toward systematic evaluation. Our analysis suggests these interpretive practices are necessary adaptations to LLM characteristics rather than methodological failures. For HCI researchers, this presents a research opportunity to support practitioners in systematizing emerging practices rather than developing new evaluation frameworks.

LLM评估人机交互实践研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。