arXiv:2601.06426cs.CLcs.AI2026-01被引 1

新基准NC-Bench从对话结构出发,评估大模型的自然对话能力。

NC-Bench: An LLM Benchmark for Evaluating Conversational Competence

  • 基于人类对话理论,聚焦对话形式而非内容
  • 六款开源模型在复杂多轮请求上表现最差
  • 适合评估对话系统的真实交互能力

自然对话基准(NC-Bench)提出一种评估大语言模型(LLMs)通用对话能力的新方法。与以往侧重模型输出内容的基准不同,NC-Bench关注自然对话的形式与结构。基于IBM自然对话框架(NCF),该基准包含三个独立数据集:(1)基础集评估基本序列管理能力,如回答问题、修复回应和结束对话对;(2)RAG集在相同模式下引入通过检索增强生成的信息获取;(3)复杂请求集扩展至涉及更复杂的序列管理模式。每个数据集测试模型对典型交互模式做出情境恰当对话行为的能力。对六款开源模型在14种交互模式上的初步评估显示,模型在基础问答任务中表现良好,但在修复任务(尤其是重复)上表现较差,关闭序列表现参差不齐,复杂多轮请求最为困难。通过操作化人类对话的基本原则,NC-Bench提供了一个轻量、可扩展且理论基础扎实的框架,用于超越主题或任务特定基准,评估和提升LLM的对话能力。

原文摘要 · Abstract (English)

The Natural Conversation Benchmark (NC-Bench) introduces a new approach to evaluating the general conversational competence of large language models (LLMs). Unlike prior benchmarks that focus on the content of model behavior, NC-Bench focuses on the form and structure of natural conversation. Grounded in the IBM Natural Conversation Framework (NCF), NC-Bench comprises three distinct sets: (1) the basic set evaluates fundamental sequence management practices, such as answering inquiries, repairing responses, and closing conversational pairs; (2) the retrieval-augmented generation (RAG) set applies the same sequence management patterns as the first set but incorporates information-seeking via RAG; (3) the complex request set extends to requests involving more intricate sequence management patterns. Each set tests a model's ability to produce contextually appropriate conversational actions in response to characteristic interaction patterns. Initial evaluations across six open-source models and 14 interaction patterns show that models perform well on basic answering tasks, struggle more with repair tasks (especially repeat), have mixed performance on closing sequences, and find complex multi-turn requests most challenging. By operationalizing fundamental principles of human conversation, NC-Bench provides a lightweight, extensible, and theory-grounded framework for assessing and improving the conversational abilities of LLMs beyond topical or task-specific benchmarks.

对话评估基准测试大模型LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。