揭示提示如何通过几何变换重塑模型内部表示以引导行为
Decomposing how prompting steers behavior

- 将提示效果分解为平移、刚性变换等几何操作,逐层验证其影响
- 发现线性混合是提示重构表示的关键机制,可显著提升行为匹配度
- 适用于研究大模型行为机制的学者,尤其关注提示工程与表征变化者
提示能在不更新权重的情况下引导大语言模型(LLMs)和视觉-语言模型(VLMs)的行为,但其如何改变内部表示仍不清楚。本文提出一种嵌套几何分解框架,将提示视为对提示后内容表征几何的转换。针对每组提示对,使用渐进表达力更强的、与刺激无关的映射(平移、带均匀缩放的刚性变换、顺序轴缩放、仿射变换、非线性变换)对同一刺激在两种提示下的表征进行对齐。随后通过因果测试:用映射后的提示A隐藏状态替换某一层中未见样本的原始状态,测量提示B表征几何与行为的恢复程度。在三个LLM、三个VLM及六个文本或图像数据集(涵盖风格、情感、场景内容、数量等)上,提示均一致地将表示重塑至目标任务结构。交叉验证方差分解显示,大量提示引起的激活变化由保持形状的映射捕获,尤其是平移和带均匀缩放的刚性变换;层级特征分析揭示了模型与任务特异的路由策略。关键发现:尽管平移与刚性层级已提升行为一致性,但仿射变换是首个几乎完全恢复目标提示任务几何并带来相应行为提升的层级,表明跨维度线性混合是提示重构表示的核心机制。该框架将提示诱导的表征变化分解为可解释的几何成分,揭示了模型如何将任务相关结构路由以生成提示驱动行为。
原文摘要 · Abstract (English)
Prompting steers large language models (LLMs) and vision-language models (VLMs) without weight updates, but it remains unclear how instruction changes reshape internal representations to produce behavior. We introduce a nested geometric decomposition framework that treats prompting as a transformation of the representational geometry of the content following the prompt. For each prompt pair, we align representations of the same stimuli under two prompts using increasingly expressive stimulus-invariant maps: translation, rigid transformation with uniform scaling, sequential axis scaling, affine transformation, and nonlinear transformation. We then causally test each map by replacing a single layer's prompt-A hidden state for held-out stimuli with its mapped counterpart and measuring recovery of prompt-B representational geometry and behavior. Across three LLMs, three VLMs, and six text or image datasets spanning style, emotion, scene content, and number, prompts consistently reshape representations toward the instructed task structure. Cross-validated variance decomposition shows that much prompt-induced activation change is captured by shape-preserving maps, especially translation and rigid transformation with uniform scaling, while tier profiles reveal model- and task-specific routing strategies across layers. Crucially, although translation and rigid tiers already improve behavioral agreement, affine transformation is the first tier to nearly recover target-prompt task geometry and yields corresponding behavioral gains. This suggests that cross-dimensional linear mixing is a key mechanism by which prompts reorganize representations toward instructed task structure. Our framework decomposes prompt-induced representational change into interpretable geometric components and reveals how models route task-relevant structure to produce prompt-driven behavior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。