让大模型自己打分,追踪对话中的情绪状态变化。
Quantitative Introspection in Language Models: Tracking Emotive States Across Conversation
- 用模型自评数值与探测器定义的状态做关联,量化其内在情绪变化。
- 自评分数能有效反映状态演变,大模型中相关性最高达0.76。
- 该方法可揭示模型认知演化,适合关注模型安全与可解释性的研究者。
追踪大语言模型在对话中的内部状态对安全性、可解释性及模型福祉至关重要,但现有方法受限。受人类心理学中数值自评启发,我们探究大模型能否通过自身数值自评来追踪探测器定义的情绪状态。在40轮十回合对话中,研究了幸福感、兴趣、专注度和冲动性四组概念,将内省操作化为模型自评与概念匹配探测状态间的因果信息耦合。发现贪婪解码的自评会坍缩为少数无信息值,而基于逻辑值的自评能揭示可解释的内部状态(斯皮尔曼ρ=0.40–0.76;等序回归R²=0.12–0.54,LLaMA-3.2-3B-Instruct),且状态随时间变化可追踪,激活操控证实耦合具有因果性。此外,内省能力从第一轮即存在并随对话演化,通过沿一个概念引导可选择性提升另一概念的内省能力(ΔR²最高达0.30)。关键的是,某些现象随模型规模增长,在LLaMA-3.1-8B-Instruct中接近R²≈0.93,且在其他模型族中部分复现。这些结果表明,数值自评是追踪对话式人工智能系统内部情绪状态的一种可行互补工具。
原文摘要 · Abstract (English)
Tracking the internal states of large language models across conversations is important for safety, interpretability, and model welfare, yet current methods are limited. Linear probes and other white-box methods compress high-dimensional representations imperfectly and are harder to apply with increasing model size. Taking inspiration from human psychology, where numeric self-report is a widely used tool for tracking internal states, we ask whether LLMs' own numeric self-reports can track probe-defined emotive states over time. We study four concept pairs (wellbeing, interest, focus, and impulsivity) in 40 ten-turn conversations, operationalizing introspection as the causal informational coupling between a model's self-report and a concept-matched probe-defined internal state. We find that greedy-decoded self-reports collapse outputs to few uninformative values, but introspective capacity can be unmasked by calculating logit-based self-reports. This metric tracks interpretable internal states (Spearman $ρ= 0.40$-$0.76$; isotonic $R^2 = 0.12$-$0.54$ in LLaMA-3.2-3B-Instruct), follows how those states change over time, and activation steering confirms the coupling is causal. Furthermore, we find that introspection is present at turn 1 but evolves through conversation, and can be selectively improved by steering along one concept to boost introspection for another ($ΔR^2$ up to $0.30$). Crucially, these phenomena scale with model size in some cases, approaching $R^2 \approx 0.93$ in LLaMA-3.1-8B-Instruct, and partially replicate in other model families. Together, these results position numeric self-report as a viable, complementary tool for tracking internal emotive states in conversational AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。