让大模型生成能化解矛盾陈述的解释,提升其推理能力。
Explanation Generation for Contradiction Reconciliation with LLMs
- 用现有自然语言推理数据集改造任务,生成可兼容矛盾陈述的解释。
- 18个大模型实验表明,多数模型在该任务上表现有限,且算力增加收益递减。
- 适合需要深度推理的对话系统与科学辅助工具开发者参考。
现有NLP研究通常将矛盾视为需选择接受或舍弃的错误,但人类在社交互动和专业领域中常通过假设解释来调和矛盾。例如,“卡西讨厌咖啡”与“她每天买咖啡”看似矛盾,若卡西每日为同事采购咖啡,则二者可共存。尽管大语言模型(LLMs)推理能力不断提升,其生成此类调和性解释的能力仍鲜有研究。为此,我们提出“调和性解释生成”任务,要求模型生成使矛盾陈述相容的解释。我们提出一种重用现有自然语言推理(NLI)数据集的方法,并引入可扩展的自动评估指标。对18个LLMs的实验表明,多数模型在此任务上成效有限,且通过“思考”扩展测试时计算量带来的收益随模型规模增大而趋于饱和。结果揭示了大模型推理的一个未被充分探索维度,凸显了提升该能力对聊天机器人、科学辅助等下游应用的重要性。
原文摘要 · Abstract (English)
Existing NLP work commonly treats contradictions as errors to be resolved by choosing which statements to accept or discard. Yet a key aspect of human reasoning in social interactions and professional domains is the ability to hypothesize explanations that reconcile contradictions. For example, "Cassie hates coffee" and "She buys coffee everyday" may appear contradictory, yet both are compatible if Cassie has the unenviable daily chore of buying coffee for all her coworkers. Despite the growing reasoning capabilities of large language models (LLMs), their ability to hypothesize such reconciliatory explanations remains largely unexplored. To address this gap, we introduce the task of reconciliatory explanation generation, where models must generate explanations that effectively render contradictory statements compatible. We propose a novel method of repurposing existing natural language inference (NLI) datasets, and introduce quality metrics that enable scalable automatic evaluation. Experiments with 18 LLMs show that most models achieve limited success in this task, and that the benefit of extending test-time compute by "thinking" plateaus as model size increases. Our results highlight an under-explored dimension of LLM reasoning and the need to address this limitation in enhancing LLMs' downstream applications such as chatbots and scientific aids.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。