arXiv:2502.19649cs.LGcs.CL2025-02被引 33

直接操控大模型内部表示,实现更灵活的可控性

Taxonomy, Opportunities, and Challenges of Representation Engineering for Large Language Models

  • 通过识别、操作和控制模型内部表征来调节行为
  • 相比微调等方法,更具可解释性和数据效率
  • 适合研究可控生成与模型干预的学者

表示工程(RepE)是一种控制大语言模型(LLM)行为的新范式。与传统修改输入或微调模型的方法不同,RepE直接操纵模型的内部表示,可能在有效性、可解释性、数据效率和灵活性方面具有优势。本文首次对面向大语言模型的RepE进行系统综述,梳理快速发展的文献,回答关键问题:存在哪些RepE方法及其差异?已应用于哪些概念与问题?与其它方法相比有何优劣?为此,我们提出一个统一框架,将RepE视为包含表示识别、操作化和控制的流程。尽管RepE潜力巨大,仍面临多概念管理、可靠性保障和性能保持等挑战。我们进一步提出实验与方法改进机会,并构建最佳实践指南。

原文摘要 · Abstract (English)

Representation Engineering (RepE) is a novel paradigm for controlling the behavior of LLMs. Unlike traditional approaches that modify inputs or fine-tune the model, RepE directly manipulates the model's internal representations. As a result, it may offer more effective, interpretable, data-efficient, and flexible control over models' behavior. We present the first comprehensive survey of RepE for LLMs, reviewing the rapidly growing literature to address key questions: What RepE methods exist and how do they differ? For what concepts and problems has RepE been applied? What are the strengths and weaknesses of RepE compared to other methods? To answer these, we propose a unified framework describing RepE as a pipeline comprising representation identification, operationalization, and control. We posit that while RepE methods offer significant potential, challenges remain, including managing multiple concepts, ensuring reliability, and preserving models' performance. Towards improving RepE, we identify opportunities for experimental and methodological improvements and construct a guide for best practices.

表示工程大模型控制可解释性方法综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。