构建复杂问答基准,评估大模型处理多步推理问题的能力。
Compound-QA: A Benchmark for Evaluating LLMs on Compound Questions
- 用合成方法构建包含多步逻辑的复合问题数据集
- 九个开源模型在复合问题上表现显著低于简单问题
- 提出改进策略,有效提升模型理解与推理能力
大语言模型在各类任务中表现优异,但现有评估基准多聚焦单个问题,忽视真实场景中的复杂交互。本文提出复合问题生成方法(CQ-Syn),构建了针对多步关联子问题的Compound-QA基准。该数据集源自现有QA数据集,由自研LLM标注并经人工验证,涵盖五类问题:事实陈述、因果关系、假设分析、比较选择与评价建议。从理解、推理、知识三个维度评估模型能力。对九个开源模型的测试显示,其在复合问题上的表现明显低于非复合问题。进一步探索优化策略后发现,这些方法显著提升了模型的综合理解与推理能力。
原文摘要 · Abstract (English)
Large language models (LLMs) demonstrate remarkable performance across various tasks, prompting researchers to develop diverse evaluation benchmarks. However, most benchmarks typically measure the ability of LLMs to respond to individual questions, neglecting the complex interactions in real-world applications. We introduce Compound Question Synthesis (CQ-Syn) to build Compound-QA, a benchmark targeting questions composed of multiple interrelated sub-questions. This benchmark is derived from existing QA datasets, annotated with proprietary LLMs, and verified by humans for accuracy. It encompasses five categories: Factual-Statement, Cause-and-Effect, Hypothetical-Analysis, Comparison-and-Selection, and Evaluation-and-Suggestion. It evaluates the LLM capability in terms of three dimensions, including understanding, reasoning, and knowledge. Evaluating nine open-source LLMs on Compound-QA reveals that their performance on compound questions is notably lower than on non-compound questions. We further explore strategies to enhance LLMs' handling of compound questions, and our results show that these methods substantially improve models' comprehension and reasoning abilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。