构建中文多轮反馈评估基准,测大模型真实交互响应能力
FB-Bench: A Fine-Grained Multi-Task Benchmark for Evaluating LLMs' Responsiveness to Human Feedback
- 设计591个细粒度样本,覆盖8类任务与9种用户反馈
- 发现不同任务/反馈类型下模型表现差异显著
- 适合研究人机交互、模型可解释性的研究人员使用
人类反馈在人与大语言模型的交互中至关重要。然而,现有研究主要聚焦单轮对话评估。即使在多轮对话基准中,用户输入也常相互独立,忽略了真实场景下反馈的复杂性。为此,我们提出FB-Bench,一个面向中文真实使用场景的细粒度多任务基准,用于评估大模型对人类反馈的响应能力。该基准包含591个精心设计的样本,涵盖8类任务、5种回复缺陷类型和9种反馈类型。我们广泛评测了多种主流大模型,发现其在不同交互场景下表现差异显著。进一步分析表明,任务类型、人类反馈形式及先前回复缺陷均显著影响模型响应。研究结果揭示了当前模型的优势与局限,为未来研究提供了重要方向。代码与数据集见https://github.com/PKU-Baichuan-MLSystemLab/FB-Bench。
原文摘要 · Abstract (English)
Human feedback is crucial in the interactions between humans and Large Language Models (LLMs). However, existing research primarily focuses on benchmarking LLMs in single-turn dialogues. Even in benchmarks designed for multi-turn dialogues, the user inputs are often independent, neglecting the nuanced and complex nature of human feedback within real-world usage scenarios. To fill this research gap, we introduce FB-Bench, a fine-grained, multi-task benchmark designed to evaluate LLMs' responsiveness to human feedback under real-world usage scenarios in Chinese. Drawing from the two main interaction scenarios, FB-Bench comprises 591 meticulously curated samples, encompassing eight task types, five deficiency types of response, and nine feedback types. We extensively evaluate a broad array of popular LLMs, revealing significant variations in their performance across different interaction scenarios. Further analysis indicates that task, human feedback, and deficiencies of previous responses can also significantly impact LLMs' responsiveness. Our findings underscore both the strengths and limitations of current models, providing valuable insights and directions for future research. Code and datasets are available at https://github.com/PKU-Baichuan-MLSystemLab/FB-Bench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。