无需人工标注,用统计方法评估大模型对话目标达成度。
Unsupervised Evaluation of Multi-Turn Objective-Driven Interactions
- 基于无标签对话数据的统计特性,设计无监督评估指标
- 可检测目标完成度与模型不确定性,不依赖理想回复
- 适用于企业级对话系统,适合大规模自动化评估
大语言模型在企业应用中广泛用于人机目标导向交互,但现有评估方法面临数据复杂、无标签、人工标注成本高、定制指标难以发现新问题、LLM评判不可靠等挑战。本文提出首个面向目标驱动交互的无监督评估体系,利用无标签交互数据的统计特性,结合微调后的LLM适应分布偏移。所提方法可实现用户目标标注、目标完成度度量及模型不确定性量化,无需依赖人类生成的理想回复。在开放域和任务特定交互数据上验证了有效性。
原文摘要 · Abstract (English)
Large language models (LLMs) have seen increasing popularity in enterprise applications where AI agents and humans engage in objective-driven interactions. However, these systems are difficult to evaluate: data may be complex and unlabeled; human annotation is often impractical at scale; custom metrics can monitor for specific errors, but not previously-undetected ones; and LLM judges can produce unreliable results. We introduce the first set of unsupervised metrics for objective-driven interactions, leveraging statistical properties of unlabeled interaction data and using fine-tuned LLMs to adapt to distributional shifts. We develop metrics for labeling user goals, measuring goal completion, and quantifying LLM uncertainty without grounding evaluations in human-generated ideal responses. Our approach is validated on open-domain and task-specific interaction data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。