稀疏自编码器擅长发现未知概念,而非处理已知问题。
Position: Use Sparse Autoencoders to Discover Unknowns
- 用稀疏自编码器挖掘模型中未被标注的潜在概念
- 在可解释性与公平性审计中识别隐藏偏差
- 适合需要探索未知模式的研究场景
尽管稀疏自编码器(SAEs)引发广泛关注,但一系列负面结果加剧了对其实用性的质疑。本文提出一个关键概念区分:即使 SAEs 在作用于已知概念时效果有限,它们在发现未知概念方面却尤为强大。这一区分调和了关于 SAEs 的矛盾观点,揭示其多种应用场景,包括机器学习可解释性、可追溯性、公平性、审计与安全,以及社会科学和健康科学中的未知模式探索。
原文摘要 · Abstract (English)
While sparse autoencoders (SAEs) have generated significant excitement, a series of negative results have added to skepticism about their usefulness. Here, we establish a conceptual distinction that reconciles competing narratives surrounding SAEs. We argue that even if SAEs may be less effective for \textit{acting on known concepts}, SAEs are especially powerful tools for \textit{discovering unknown concepts}. This distinction separates existing negative results from positive results, and suggests several classes of SAE applications. Specifically, we outline use cases for SAEs in (i) ML interpretability, explainability, fairness, auditing, and safety, and (ii) social and health sciences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。