通过对比样本检测与编辑概念表征,提升大模型的可预测性与安全性。
Representation Engineering for Large-Language Models: Survey and Research Challenges
- 用对比输入样本来定位和修改模型中的高阶概念表征
- 能有效干预如诚实、有害性等抽象概念的输出行为
- 适合关注模型安全、可控性和个性化定制的研究者
大语言模型虽能完成多种任务,但仍存在不可预测和难控制的问题。表示工程通过使用具有对比性的输入样本,来探测并编辑诸如诚实性、危害性或权力追求等高层次概念的内部表征。本文系统梳理了表示工程的目标与方法,构建了该新兴领域的统一图景。相比机制可解释性、提示工程和微调等方法,表示工程更专注于概念层面的干预。同时指出潜在风险,包括性能下降、计算时间增加及可控性问题。最后提出未来研究议程,旨在实现可预测、动态、安全且可个性化的大型语言模型。
原文摘要 · Abstract (English)
Large-language models are capable of completing a variety of tasks, but remain unpredictable and intractable. Representation engineering seeks to resolve this problem through a new approach utilizing samples of contrasting inputs to detect and edit high-level representations of concepts such as honesty, harmfulness or power-seeking. We formalize the goals and methods of representation engineering to present a cohesive picture of work in this emerging discipline. We compare it with alternative approaches, such as mechanistic interpretability, prompt-engineering and fine-tuning. We outline risks such as performance decrease, compute time increases and steerability issues. We present a clear agenda for future research to build predictable, dynamic, safe and personalizable LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。