arXiv:2607.00083cs.CLcs.AI2026-07

通过潜空间控制与校准,提升大模型行为可控性与可信度。

Harnessing the Latent Space: From Steering Vectors to Model Calibrators for Control and Trust

论文配图:Harnessing the Latent Space: From Steering Vectors to Model Calibrators for Control and Trust
图 1 · 摘自论文原文
  • 提出方向向量实现对语言模型输出的精准控制。
  • 开发基于潜空间的校准器,判断模型输出可信度。
  • 适合关注模型可解释性与安全性的研究者与工程师。

语言模型已从不可靠的文本生成器演变为参数量达万亿级的强大模型。随着模型规模扩大,理解其内部表示愈发困难。由于数百万用户在中高风险场景中依赖语言模型与外部工具交互或做出决策,建立对模型行为的控制机制并判断其输出可信度变得至关重要。本文提出利用潜空间的方法:一是设计方向向量以实现对模型行为的操控;二是开发基于潜空间的模型校准器,用于评估输出可靠性。这两项工作共同揭示了语言模型潜空间的内在规律,为构建更可信的语言技术提供了新思路。

原文摘要 · Abstract (English)

Language models have changed from unreliable text generators to highly-capable large models with trillions of parameters. Capability increases come hand-in-hand with increases in scale, making understanding the internal representations of models more challenging. Since millions of users increasing rely on language models to interact with external tools or make decisions in medium or high-stakes scenarios, we need to establish control over model behavior and know when to trust model outputs. In this paper, we discuss our contributions on harnessing the latent spaces by proposing steering vectors for control and developing latent space-based model calibrators for trust. Together, our contributions help demystify the latent spaces of language models and offer new insights into how to harness model internals to build more trustworthy language technology.

潜空间模型控制可信AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。