arXiv:2502.08009cs.CL2025-02NAACL被引 14

揭示提示如何通过不同表征机制让大模型灵活切换任务

The Geometry of Prompting: Unveiling Distinct Mechanisms of Task Adaptation in Language Models

  • 用统计物理框架分析提示对模型表征几何的影响
  • 不同提示方法性能相似但内在机制迥异
  • 发现任务间存在表征层面的协同与干扰

仅解码器架构的语言模型能根据输入提示动态切换计算任务。尽管提示技术应用广泛,其内部灵活性机制仍不明确。本文基于统计物理框架,研究不同提示方法如何影响模型表征几何。结果表明,虽各类提示达成相似性能,却通过截然不同的表征机制实现任务适配。分析揭示输入分布样本和标签语义在少样本上下文学习中的关键作用,并首次提供任务间表征层面协同与干扰的实证证据。本工作深化了对大模型理论的理解,为开发更高效、表征感知的提示策略奠定基础。

原文摘要 · Abstract (English)

Decoder-only language models have the ability to dynamically switch between various computational tasks based on input prompts. Despite many successful applications of prompting, there is very limited understanding of the internal mechanism behind such flexibility. In this work, we investigate how different prompting methods affect the geometry of representations in these models. Employing a framework grounded in statistical physics, we reveal that various prompting techniques, while achieving similar performance, operate through distinct representational mechanisms for task adaptation. Our analysis highlights the critical role of input distribution samples and label semantics in few-shot in-context learning. We also demonstrate evidence of synergistic and interfering interactions between different tasks on the representational level. Our work contributes to the theoretical understanding of large language models and lays the groundwork for developing more effective, representation-aware prompting strategies.

语言模型提示工程表征几何

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。