让AI概念可视觉探索,提升大模型内部理解力
ConceptViz: A Visual Analytics Approach for Exploring Concepts in Large Language Models
- 构建识别-解释-验证三步流程,连接稀疏自编码器特征与人类概念
- 用户可通过概念查询、交互探索特征对齐关系,验证模型行为
- 适合研究大模型可解释性、认知机制的学者和工程师
大语言模型在自然语言任务中表现卓越,但其内部知识表征机制仍不清晰。尽管稀疏自编码器(SAEs)是提取可解释特征的有前景方法,但其特征与人类可理解概念之间缺乏天然对应,导致解释过程繁琐耗时。为此,我们提出ConceptViz——一个用于探索大语言模型中概念的可视化分析系统。该系统采用新型的‘识别→解释→验证’流程,支持用户以感兴趣的概念为查询,交互式探索概念与特征之间的对齐关系,并通过模型行为验证对应关系的合理性。我们在两个使用场景和一次用户研究中验证了ConceptViz的有效性。结果表明,该系统显著提升了概念表征发现与验证的效率,帮助研究者建立更准确的大模型特征认知模型。代码与用户指南已开源:https://github.com/Happy-Hippo209/ConceptViz。
原文摘要 · Abstract (English)
Large language models (LLMs) have achieved remarkable performance across a wide range of natural language tasks. Understanding how LLMs internally represent knowledge remains a significant challenge. Despite Sparse Autoencoders (SAEs) have emerged as a promising technique for extracting interpretable features from LLMs, SAE features do not inherently align with human-understandable concepts, making their interpretation cumbersome and labor-intensive. To bridge the gap between SAE features and human concepts, we present ConceptViz, a visual analytics system designed for exploring concepts in LLMs. ConceptViz implements a novel dentification => Interpretation => Validation pipeline, enabling users to query SAEs using concepts of interest, interactively explore concept-to-feature alignments, and validate the correspondences through model behavior verification. We demonstrate the effectiveness of ConceptViz through two usage scenarios and a user study. Our results show that ConceptViz enhances interpretability research by streamlining the discovery and validation of meaningful concept representations in LLMs, ultimately aiding researchers in building more accurate mental models of LLM features. Our code and user guide are publicly available at https://github.com/Happy-Hippo209/ConceptViz.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。