arXiv:2502.05242cs.CLcs.AI2025-02被引 2

让大模型自己更容易被监控,提升透明度和安全性。

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring

  • 通过内嵌机制增强大模型思维过程的可观察性
  • 在多模态测试集上实现一致性能提升,适配不同架构与参数规模
  • 结合最优传输理论分析泛化能力改进,适合安全与伦理研究者

大语言模型(LLMs)能力日益增强,但其思考与决策机制仍不清晰。链式思维(CoTs)虽常用于外显模型推理,却无法准确反映真实思维过程。基于隐藏表示的方法提供了内部视角,但以往工作仅依赖外部模块,未真正让模型自身更易监控。本文提出新方法TELLME,提升LLM透明度,帮助识别不当与敏感行为。实验显示,TELLME在去毒任务中对多模态测试集、不同架构及不同参数规模的模型均实现一致改进。进一步从最优传输理论与实证角度分析了TELLME对模型泛化能力的提升效果。

原文摘要 · Abstract (English)

Large language models (LLMs) are becoming increasingly capable, but the mechanisms of their thinking and decision-making processes remain unclear. Chain-of-thoughts (CoTs) have been commonly utilized to externalize LLMs' thinking, but this strategy fails to accurately reflect LLMs' thinking process. Techniques based on LLMs' hidden representations provide an inner perspective to improve the monitorability of their latent thinking. However, previous methods only try to develop external modules instead of making LLMs themselves easier to monitor. In this paper, we propose a novel method, TELLME, improving the transparency of LLMs and helping monitors identify unsuitable and sensitive behaviors. Furthermore, we showcase the effectiveness of TELLME on detoxification tasks, where LLMs achieve consistent improvement among multimodal test sets, distinct architectures, and varying parameter scales. We further analyze TELLME's improvement on LLMs' generalization ability from both optimal transport theory and empirical perspectives.

大模型监控透明性安全对齐泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。