arXiv:2607.18259cs.AI2026-07

让大模型生成更可信,通过概率化控制概念偏向。

Probabilistic Concept-Aware Steering for Trustworthy LLM Inference

论文配图:Probabilistic Concept-Aware Steering for Trustworthy LLM Inference
图 1 · 摘自论文原文
  • 基于概念检索与概率校准,动态生成导向向量
  • 在保持原任务能力下实现可控语义偏移
  • 适合需要安全可控生成的场景,如医疗、金融

转向向量(SVs)是一种大语言模型推理阶段的干预技术,通过在推理过程中向中间激活添加特定概念的方向向量来引导生成。然而,现有方法常导致表征不一致,影响可解释性与细粒度控制,主要原因在于以往研究仅关注正负二元评价,且使用离散聚类指标,无法捕捉语义对齐的连续谱。本文提出概率化概念感知转向框架(PCS),在保持原始任务能力的同时,通过概念驱动的转向向量检索与概率强度校准,实现可控、面向安全的语义偏置。

原文摘要 · Abstract (English)

Steering vectors (SVs), an inference-time intervention technique for large language models (LLMs), guide the generation process by adding a concept-specific direction vector to intermediate activations during inference. However, existing SV methods frequently yield representation-incoherent behaviors that undermine interpretability and fine-grained control, largely because prior work has focused on binary positive-negative steering evaluation while employing discrete clustering metrics that fail to capture the continuous spectrum of semantic alignment. In this work, we present the Probabilistic Concept-Aware Steering (PCS) framework for LLM inference. PCS preserves original task competence while providing controllable, safety-oriented semantic bias through concept-driven steering-vector retrieval and probabilistic strength calibration.

大模型推理概念控制可信生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。