用跨领域通用逻辑结构提升大模型推理能力
Towards Effective In-context Cross-domain Knowledge Transfer via Domain-invariant-neurons-based Retrieval
- 基于领域不变神经元提取通用逻辑表示,实现跨域示例检索
- 在数学与逻辑推理任务中平均性能超越现有方法1.8分
- 适合缺乏专家标注的冷门领域如法律、形式逻辑
大语言模型在逻辑推理方面已取得显著进展,但仍远未达到人类水平。当前增强策略依赖专家手工构建的本域示例,限制了在专业性稀缺领域的应用,如专门数学推理、形式逻辑或法律分析。本文证明了利用跨域示例提升大模型推理性能的可行性。尽管领域差异显著,但许多可复用的隐式逻辑结构在不同领域间共享。为此,我们提出一种有效的检索方法——基于领域不变神经元的检索(DIN-Retrieval)。该方法首先提取跨领域通用的隐藏表示,推理时使用DIN向量检索结构兼容的跨域示例用于上下文学习。在多个数学与逻辑推理任务上的迁移实验表明,该方法平均性能比当前最优方法提升1.8分。
原文摘要 · Abstract (English)
Large language models (LLMs) have made notable progress in logical reasoning, yet still fall short of human-level performance. Current boosting strategies rely on expert-crafted in-domain demonstrations, limiting their applicability in expertise-scarce domains, such as specialized mathematical reasoning, formal logic, or legal analysis. In this work, we demonstrate the feasibility of leveraging cross-domain demonstrating examples to boost the LLMs' reasoning performance. Despite substantial domain differences, many reusable implicit logical structures are shared across domains. In order to effectively retrieve cross-domain examples for unseen domains under investigation, in this work, we further propose an effective retrieval method, called domain-invariant neurons-based retrieval (\textbf{DIN-Retrieval}). Concisely, DIN-Retrieval first summarizes a hidden representation that is universal across different domains. Then, during the inference stage, we use the DIN vector to retrieve structurally compatible cross-domain demonstrations for the in-context learning. Experimental results in multiple settings for the transfer of mathematical and logical reasoning demonstrate that our method achieves an average improvement of 1.8 over the state-of-the-art methods \footnote{Our implementation is available at https://github.com/Leon221220/DIN-Retrieval}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。