arXiv:2606.24834cs.AI2026-06中稿 · SIGDIAL 2026

研究大模型对话在非功能需求评估中的准确性和用户体验

Accuracy and Satisfaction in Multi-Turn LLM Dialogues for NFR Assessment

论文配图:Accuracy and Satisfaction in Multi-Turn LLM Dialogues for NFR Assessment
图 1 · 摘自论文原文
  • 构建多轮对话评估框架,分析开发者与大模型协作评估合规性需求
  • 大模型输出与专家判断一致性低,但开发者普遍认可其结论
  • 主动交互提升满意度,冗长回复反而降低体验

基于大模型的对话助手已成为开发者的主流工具,但现有评估基准仅关注功能正确性,忽略了对非功能需求(NFRs)评估中对话质量与准确性的重要性。本文聚焦医疗数据隐私法规HIPAA合规场景,招募49名程序员通过多轮对话使用GitHub Copilot评估iTrust代码库中的148个基于HIPAA的NFR,从需求满足度、推理过程和代码定位三个维度进行分析。结果发现,开发者对大模型的判断普遍认同,但与专家基准相比准确性较低。研究还发现,系统响应越长、信息提供越多的回合会降低用户满意度,而主动引导的互动则显著提升满意度。这些发现为设计支持NFR评估的对话系统提供了实证依据。

原文摘要 · Abstract (English)

LLM-based dialogue assistants have become mainstream tools for software developers, yet current evaluation benchmarks focus exclusively on functional correctness. This leaves a critical gap in assessing the quality and accuracy of these conversations when handling Non-Functional Requirements (NFRs), which are inherently vague, context-dependent, and involve many parts of a program. Evaluating how well these systems support collaborative reasoning about NFRs requires methods that go beyond single-turn accuracy to capture both the correctness of the system's outputs and the quality of the multi-turn interaction. In this paper, we investigate the accuracy and quality of multi-turn conversations between developers and an LLM-based agent in the domain of Health Insurance Portability and Accountability Act (HIPAA) regulatory compliance. We hired 49 programmers to interact with GitHub Copilot to assess 148 HIPAA-derived NFRs against the iTrust codebase, a system designed to comply with HIPAA regulations, across three dimensions: requirement satisfaction level, reasoning, and code localization. We find that developers tend to agree with LLM assessments, but accuracy against expert ground truth is low. We model user satisfaction and find that longer system responses and more information-providing turns negatively affect user satisfaction, whereas proactive interactions positively affect it. Our findings provide insights for designing LLM-based dialogue systems that support NFR assessment.

大模型对话非功能需求合规评估用户满意度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。