让AI辩论更真实可信,从情感到逻辑全面评估并优化。
InspireDebate: Multi-Dimensional Subjective-Objective Evaluation-Guided Reasoning and Optimization for Debating
- 设计多维度评估体系,融合主观感受与客观事实检验。
- 相比基线模型,辩论表现提升57%,专家评分相关性提高44%。
- 适合需要高可信度辩论生成的场景,如教育、政策分析。
随着大语言模型的发展,辩论任务如论点质量评估和辩论过程模拟取得了显著进展。然而,现有基于LLM的辩论系统多聚焦于回应特定论点,忽视了真实性与逻辑有效性等客观评估。此外,这些系统缺乏在评估指标、思维链推理及多轮辩论优化等多个维度上的结构化优化方法,制约了其效果。为此,我们提出双组件框架:(1) InspireScore,一种包含四大主观标准(情感吸引力、论点清晰度、论证结构、话题相关性)和两大客观指标(事实真实性、逻辑有效性)的多维评估架构;(2) InspireDebate,通过思维链增强、多维直接偏好优化(DPO)以及基于网络的检索增强生成(Web-RAG)实现分阶段优化。实证评估表明,InspireScore与专家判断的相关性比现有方法高出44%,InspireDebate相较基线模型性能提升57%。源代码已开源:https://github.com/fywang12/InspireDebate。
原文摘要 · Abstract (English)
With the rapid advancements in large language models (LLMs), debating tasks, such as argument quality assessment and debate process simulation, have made significant progress. However, existing LLM-based debating systems focus on responding to specific arguments while neglecting objective assessments such as authenticity and logical validity. Furthermore, these systems lack a structured approach to optimize across various dimensions$-$including evaluation metrics, chain-of-thought (CoT) reasoning, and multi-turn debate refinement$-$thereby limiting their effectiveness. To address these interconnected challenges, we propose a dual-component framework: (1) $\textbf{InspireScore}$, a novel evaluation system that establishes a multi-dimensional assessment architecture incorporating four subjective criteria (emotional appeal, argument clarity, argument arrangement, and topic relevance) alongside two objective metrics (fact authenticity and logical validity); and (2) $\textbf{InspireDebate}$, an optimized debating framework employing a phased optimization approach through CoT reasoning enhancement, multi-dimensional Direct Preference Optimization (DPO), and real-time knowledge grounding via web-based Retrieval Augmented Generation (Web-RAG). Empirical evaluations demonstrate that $\textbf{InspireScore}$ achieves 44$\%$ higher correlation with expert judgments compared to existing methods, while $\textbf{InspireDebate}$ shows significant improvements, outperforming baseline models by 57$\%$. Source code is available at https://github.com/fywang12/InspireDebate.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。