让普通用户看清AI行为背后的神经机制,提前预判聊天机器人表现。
Neural Transparency: Mechanistic Interpretability Interfaces for Anticipating Model Behaviors for Personalized AI
- 通过对比提示词激活差异,提取情感、讨好、毒性等行为向量。
- 用户对15种特质中11种判断错误,暴露了认知偏差。
- 交互式太阳图可视化提升信任感,适合非技术用户使用。
数百万用户正在设计个性化的基于大语言模型的聊天机器人,但难以准确预判其部署后的实际行为。这种不透明性导致看似无害的提示可能引发过度讨好、毒性等不良表现,影响实用性并带来安全风险。为此,我们提出一种神经透明接口,在聊天机器人设计阶段揭示模型内部机制。该方法通过计算对比性系统提示词在神经激活上的差异,提取情感、毒性、讨好等行为特征向量,并将系统提示的最终词元激活投影到这些向量上,进行归一化后以交互式太阳图呈现。我们在Prolific上开展在线用户研究,对比本接口与无透明度基线界面。分析显示,参与者对15种可分析特质中有11种存在系统性误判,凸显透明工具的必要性。尽管界面未改变设计迭代模式,但显著提升了用户信任度,反馈积极。定性分析揭示了用户对可视化界面的复杂体验,为未来优化提供方向。本工作为机制可解释性面向非技术用户的落地提供了路径,奠定了更安全、对齐的人机交互基础。
原文摘要 · Abstract (English)
Millions of users now design personalized LLM-based chatbots that shape their daily interactions, yet they can only roughly anticipate how their design choices will manifest as behaviors in deployment. This opacity is consequential: seemingly innocuous prompts can trigger excessive sycophancy, toxicity, or other undesirable traits, degrading utility and raising safety concerns. To address this issue, we introduce an interface that enables neural transparency by exposing language model internals during chatbot design. Our approach extracts behavioral trait vectors (empathy, toxicity, sycophancy, etc.) by computing differences in neural activations between contrastive system prompts that elicit opposing behaviors. We predict chatbot behaviors by projecting the system prompt's final token activations onto these trait vectors, normalizing for cross-trait comparability, and visualizing results via an interactive sunburst diagram. To evaluate this approach, we conducted an online user study using Prolific to compare our neural transparency interface against a baseline chatbot interface without any form of transparency. Our analyses suggest that users systematically miscalibrated AI behavior: participants misjudged trait activations for eleven of fifteen analyzable traits, motivating the need for transparency tools in everyday human-AI interaction. While our interface did not change design iteration patterns, it significantly increased user trust and was enthusiastically received. Qualitative analysis revealed nuanced user experiences with the visualization, suggesting interface and interaction improvements for future work. This work offers a path for how mechanistic interpretability can be operationalized for non-technical users, establishing a foundation for safer, more aligned human-AI interactions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。