arXiv:2504.15630cs.CL2025-04ACL被引 3

通过增强模型内部表示,让大模型更好理解上下文知识。

Exploiting Contextual Knowledge in LLMs through V-usable Information based Layer Enhancement

  • 基于可利用信息分析,在合适层增强上下文信息
  • 问答任务中显著提升未知或冲突上下文下的生成准确性
  • 适合关注模型推理可靠性与上下文理解的研究者

大型语言模型在多项任务中表现出色,但在生成需忠实反映上下文知识的内容时仍存在不足。现有方法多聚焦于解码策略优化,却忽视了上下文信息在模型内部状态中的处理机制。为此,本文提出上下文感知层增强(CaLE),一种新型干预方法,通过V-usable信息分析,在最优层战略性地增强上下文信息的增长,从而丰富最终层的表征。实验表明,CaLE能有效提升问答任务中对未知或矛盾上下文的知识忠实生成能力。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have demonstrated remarkable capabilities in various tasks, yet they often struggle with context-faithfulness generations that properly reflect contextual knowledge. While existing approaches focus on enhancing the decoding strategies, they ignore the fundamental mechanism of how contextual information is processed within LLMs' internal states. As a result, LLMs remain limited in their ability to fully leverage contextual knowledge. In this paper, we propose Context-aware Layer Enhancement (CaLE), a novel intervention method that enhances the utilization of contextual knowledge within LLMs' internal representations. By employing V-usable information analysis, CaLE strategically amplifies the growth of contextual information at an optimal layer, thereby enriching representations in the final layer. Our experiments demonstrate that CaLE effectively improves context-faithful generation in Question-Answering tasks, particularly in scenarios involving unknown or conflicting contextual knowledge.

大模型上下文理解表征增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。