用AI自动生成解释,让非专家也能看懂大模型分析结果。
Simplifying Outcomes of Language Model Component Analyses with ELIA
- 用视觉语言模型自动把复杂图表转成自然语言说明。
- 用户研究显示非专家理解力显著提升,经验不影响效果。
- 交互式界面比静态图更有效,适合想了解模型的普通用户。
尽管机制可解释性已发展出强大工具来分析大语言模型(LLMs)内部运作,但其复杂性造成了使用门槛,限制了非专业人士的应用。我们设计并评估了ELIA(可解释语言可解释性分析),一个交互式网页应用,旨在简化多种语言模型组件分析的结果呈现,以服务更广泛受众。系统整合了归因分析、功能向量分析和电路追踪三种技术,并引入新方法:利用视觉语言模型自动生成自然语言解释(NLEs),用于解释这些方法产生的复杂可视化结果。通过混合方法用户研究验证了该方法的有效性,结果显示用户更偏好交互式、可探索的界面而非简单静态可视化。关键发现是,AI生成的解释有助于缩小知识差距;统计分析表明用户先前的LLM经验与理解得分之间无显著相关性,说明系统降低了不同经验水平用户的理解障碍。结论是,AI确实能简化复杂模型分析,但其真正潜力需结合注重互动性、精确性和叙事引导的以用户为中心的设计才能释放。
原文摘要 · Abstract (English)
While mechanistic interpretability has developed powerful tools to analyze the internal workings of Large Language Models (LLMs), their complexity has created an accessibility gap, limiting their use to specialists. We address this challenge by designing, building, and evaluating ELIA (Explainable Language Interpretability Analysis), an interactive web application that simplifies the outcomes of various language model component analyses for a broader audience. The system integrates three key techniques -- Attribution Analysis, Function Vector Analysis, and Circuit Tracing -- and introduces a novel methodology: using a vision-language model to automatically generate natural language explanations (NLEs) for the complex visualizations produced by these methods. The effectiveness of this approach was empirically validated through a mixed-methods user study, which revealed a clear preference for interactive, explorable interfaces over simpler, static visualizations. A key finding was that the AI-powered explanations helped bridge the knowledge gap for non-experts; a statistical analysis showed no significant correlation between a user's prior LLM experience and their comprehension scores, suggesting that the system reduced barriers to comprehension across experience levels. We conclude that an AI system can indeed simplify complex model analyses, but its true power is unlocked when paired with thoughtful, user-centered design that prioritizes interactivity, specificity, and narrative guidance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。