arXiv:2602.23610cs.CLcs.AI2026-02被引 1

用大模型生成真实场景下的多轮任务对话,提升推理评估可信度

LLM-Driven Multi-Turn Task-Oriented Dialogue Synthesis for Realistic Reasoning

  • 用三重优化框架自动生成贴近真实任务的多轮对话
  • 合成数据使模型推理挑战显著增强,推动能力提升
  • 适合研究真实推理能力评估与提升的学者使用

大型语言模型的推理能力——基于输入信息进行分析、推断和决策的能力——对构建智能任务导向对话系统至关重要。然而,现有基准未能充分反映现实场景的复杂性,限制了其在实际情境中评估和提升模型推理能力的效果。许多现有推理数据集过于简单抽象,脱离真实任务流程、领域约束和操作规则,难以有效评估模型逻辑推理能力。此外,预训练语料中的数据污染降低了评估结果可靠性,传统众包方式构建数据集又耗时且难扩展。为此,我们提出一种基于大模型的框架,用于生成扎根于真实推理场景的多轮任务导向对话,通过三重优化提升对话质量。该方法生成的对话基于真实任务场景,融入现实信息,并具备强上下文连贯性。围绕这些对话设计的推理任务被迭代优化,持续提升任务质量与挑战性。生成的数据集可作为评估和推进大模型真实逻辑推理能力的重要基准。实验表明,基于合成数据的推理任务引入了非平凡的推理挑战,有效支持了大模型推理能力的提升。

原文摘要 · Abstract (English)

The reasoning capability of large language models (LLMs), defined as their ability to analyze, infer, and make decisions based on input information, is essential for building intelligent task-oriented dialogue systems. However, existing benchmarks do not sufficiently reflect the complexity of real-world scenarios, which limits their effectiveness in evaluating and enhancing LLM reasoning in practical contexts. Many current reasoning datasets are overly simplistic and abstract, often disconnected from realistic task flows, domain constraints, and operational rules, making it difficult to effectively evaluate LLMs' logical reasoning ability. In addition, data contamination from pretraining corpora undermines the reliability of evaluation results, and traditional crowdsourcing methods for dataset construction are labor-intensive and difficult to scale. To address these challenges, we propose a LLM-driven framework for synthesizing multi-turn, task-oriented dialogues grounded in realistic reasoning scenarios, leveraging trilevel optimization to enhance dialogue quality. Our method generates dialogues grounded in authentic task scenarios, enriched with real-world information, and exhibiting strong contextual coherence. Corresponding reasoning tasks are carefully designed around these dialogues and iteratively refined to continuously improve the tasks' quality and challenge. The resulting dataset serves as a valuable benchmark for assessing and advancing the realistic logical reasoning capabilities of LLMs. Experimental results show that our synthetic data-based reasoning tasks introduce non-trivial reasoning challenges and provide meaningful support for improving the reasoning capabilities of LLMs.

推理评估对话生成大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。