arXiv:2505.06262cs.LGcs.AI2025-05ACL被引 5

Dialz让调整大模型概念输出更简单,支持快速实验与分析。

Dialz: A Python Toolkit for Steering Vectors

  • 模块化工具包,可直接操作模型激活值来增强或减弱特定概念
  • 能有效减少刻板印象等有害输出,揭示模型各层行为差异
  • 适合关注可控生成、模型可解释性的研究人员和开发者

我们介绍Dialz,一个用于开源大语言模型控制向量研究的Python框架。控制向量允许用户在推理时修改模型激活值,以放大或削弱如诚实性、积极度等概念,相比提示工程或微调更具灵活性。Dialz支持创建对比数据集、计算与应用控制向量、可视化分析等多样化任务。其强调模块化与易用性,既支持快速原型设计,也支持深入分析。实验表明,Dialz可有效降低模型输出中的刻板印象等有害内容,并揭示模型在不同层的行为特征。项目已开源,配套完整文档、教程及对主流开源模型的支持,旨在推动安全可控语言生成的研究。Dialz加速研究迭代,助力提升模型可解释性,为更安全、透明、可靠的AI系统铺路。

原文摘要 · Abstract (English)

We introduce Dialz, a framework for advancing research on steering vectors for open-source LLMs, implemented in Python. Steering vectors allow users to modify activations at inference time to amplify or weaken a 'concept', e.g. honesty or positivity, providing a more powerful alternative to prompting or fine-tuning. Dialz supports a diverse set of tasks, including creating contrastive pair datasets, computing and applying steering vectors, and visualizations. Unlike existing libraries, Dialz emphasizes modularity and usability, enabling both rapid prototyping and in-depth analysis. We demonstrate how Dialz can be used to reduce harmful outputs such as stereotypes, while also providing insights into model behaviour across different layers. We release Dialz with full documentation, tutorials, and support for popular open-source models to encourage further research in safe and controllable language generation. Dialz enables faster research cycles and facilitates insights into model interpretability, paving the way for safer, more transparent, and more reliable AI systems.

控制向量可解释性大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。