通过强化隐式推理能力,提升大模型对复杂指令的理解与执行效果。
ImpRIF: Stronger Implicit Reasoning Leads to Better Complex Instruction Following
- 将复杂指令建模为可验证的推理图,实现程序化验证与图驱动思维链。
- 在5个基准上显著超越基线模型,证明隐式推理增强有效提升指令遵循能力。
- 适合需要精准理解复杂逻辑指令的场景,如智能助手、自动化流程设计。
随着大语言模型应用日趋复杂,对鲁棒复杂指令遵循能力的需求日益增长。我们认为,深入理解指令本身,特别是其隐含的推理结构,是提升指令遵循能力的关键。因此,我们聚焦于涉及隐式推理、复杂逻辑关系及多约束依赖的复杂指令。提出ImpRIF方法,增强大模型对隐式推理指令的理解,从而提升其遵循复杂指令的能力。我们将此类指令形式化为可验证的推理图,实现程序化验证与图驱动的思维链推理。基于此,我们合成大规模单轮与多轮数据,提出基于图推理的微调方法,并采用强化学习显式训练模型沿图进行推理。在五个复杂指令遵循基准上,我们的模型显著优于基线模型。结果表明,增强隐式推理能力可显著提升复杂指令遵循性能。
原文摘要 · Abstract (English)
As applications of large language models (LLMs) become increasingly complex, the demand for robust complex instruction following capabilities is growing accordingly. We argue that a thorough understanding of the instruction itself, especially the latent reasoning structure embedded between the lines, is crucial for improving instruction following. Therefore we target complex instructions that involve implicit reasoning, intricate logical relations, and multi-constraint dependencies. We propose ImpRIF, a method to enhance LLMs' understanding of implicit reasoning instructions, thereby improving its ability to follow complex instructions. We formalize such instructions as verifiable reasoning graphs, enabling programmatic verification and graph-driven chain-of-thought reasoning. Based on this formulation, we synthesize large-scale single- and multi-turn data, propose fine-tuning with graph reasoning, and apply reinforcement learning to explicitly train models to reason along the graph. On five complex instruction following benchmarks, our models substantially outperform their base models. These results demonstrate that enhancing implicit reasoning capabilities can significantly improve complex instruction following.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。