用结构化语言让大模型特征描述更准确一致
Semantic Regexes: Auto-Interpreting LLM Features with a Structured Language
- 设计可组合的语义正则表达式,统一描述模型特征
- 在多个测试中准确度与自然语言相当,但更简洁一致
- 适合需要理解模型内部机制的研究者和开发者
自动化可解释性旨在将大语言模型(LLM)的特征转化为人类可理解的描述。然而,自然语言描述常模糊、不一致且需人工重命名。为此,我们提出语义正则表达式,一种结构化的特征描述语言。通过结合捕捉语言与语义模式的基元,以及用于上下文、组合和量化的修饰符,语义正则表达式能生成精确且丰富的特征描述。在定量基准与定性分析中,其准确度与自然语言相当,但描述更简洁、一致。其内在结构支持新型分析,包括跨层量化特征复杂度,实现从单个特征洞察到模型整体模式的可扩展解释。最后,用户研究表明,语义正则表达式有助于人们构建对模型特征的准确心智模型。
原文摘要 · Abstract (English)
Automated interpretability aims to translate large language model (LLM) features into human understandable descriptions. However, natural language feature descriptions can be vague, inconsistent, and require manual relabeling. In response, we introduce semantic regexes, structured language descriptions of LLM features. By combining primitives that capture linguistic and semantic patterns with modifiers for contextualization, composition, and quantification, semantic regexes produce precise and expressive feature descriptions. Across quantitative benchmarks and qualitative analyses, semantic regexes match the accuracy of natural language while yielding more concise and consistent feature descriptions. Their inherent structure affords new types of analyses, including quantifying feature complexity across layers, scaling automated interpretability from insights into individual features to model-wide patterns. Finally, in user studies, we find that semantic regexes help people build accurate mental models of LLM features.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。