arXiv:2605.09159cs.AI2026-05

通过动态追踪角色向量,揭示大模型推理过程中的内在对话。

Do LLMs Experience an Internal Polylogue? Investigating Reasoning through the Lens of Personas

论文配图:Do LLMs Experience an Internal Polylogue? Investigating Reasoning through the Lens of Personas
图 1 · 摘自论文原文
  • 将角色向量视为推理时的动态信号,监测其与隐藏激活的对齐变化。
  • 发现该动态信号可预测生成正确性,且效果接近低维激活总结。
  • 提出分阶段干预策略,适合想优化模型推理过程的研究者。

近期研究表明,大语言模型(LLMs)在激活空间中以线性方向编码行为特征(即“角色向量”),以往工作将其作为静态控制手段。本文将其视为动态信号:在推理过程中持续监测并干预。我们提出“多言”(polylogue)概念,指角色向量与隐藏激活随时间的对齐序列。在四个开源模型上实验显示,多言特征包含与低维激活摘要相当的正确性预测能力,同时仍可通过对应的角色方向解释。该信号还提供了具体的干预目标——在响应不同阶段调节哪些潜在方向。我们实现了一种基于段落条件的干预方法,在三个模型上提升了准确率,但第四模型性能下降,表明阶段感知的潜在空间调控可行但尚未稳定。总体而言,多言为推理过程中的监控与干预提供了一个可解释工具。

原文摘要 · Abstract (English)

Recent work shows that large language models (LLMs) encode behavioral traits ("personas") as linear directions in activation space, often called "persona vectors". Prior work has used such directions as static handles for behavioral steering. We instead treat them as dynamic signals: probes we can monitor and intervene on as reasoning unfolds. We use the term polylogue to denote the time series of alignments between persona vectors and hidden activations over the course of generation. Experiments across four open-weight models show that polylogue features contain predictive signal for correctness comparable to low-dimensional activation summaries, while remaining interpretable through their associated persona directions. They also suggest concrete steering targets, namely, which latent directions to modulate at different stages of a response. We instantiate this as a simple paragraph-conditioned intervention that improves accuracy on three of the four models but degrades the fourth, suggesting that stage-aware latent steering is possible but not yet robust. Together, this positions the polylogue as an interpretable tool for reasoning-time monitoring and intervention.

大模型推理角色向量可解释性动态干预

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。