arXiv:2508.06963cs.AIcs.LG2025-08被引 1

用多智能体自动生成并动态调整提示,让大模型更可信且无需重新训练。

MASteer: Multi-Agent Adaptive Steer Strategy for End-to-End LLM Trustworthiness Repair

  • 通过多智能体自动生成高质量修复样本,无需人工设计。
  • 在两个大模型上分别提升15.36%和4.21%的可信度指标。
  • 适合需要快速、灵活修复模型行为的开发者与部署场景。

大语言模型存在持续且演进的可信度问题,亟需自动化、灵活的修复方法以支持多样场景部署。现有方法如监督微调(SFT)和基于人类反馈的强化学习(RLHF)成本高、速度慢,而提示工程缺乏鲁棒性和可扩展性。表示工程通过推理时注入目标概念向量实现轻量化、免训练的控制,但当前方法依赖人工设计样本和固定策略,限制了自动化与适应性。为此,我们提出首个基于表示工程的端到端可信度修复框架MASteer。该框架集成两大核心组件:AutoTester,一个生成多样化、高质量引导样本的多智能体系统;以及AutoRepairer,通过锚点向量构建自适应策略,在推理时实现上下文感知的自动选择。在标准与定制化可信度任务上的实验表明,MASteer持续优于基线,分别在LLaMA-3.1-8B-Chat和Qwen-3-8B-Chat上提升15.36%和4.21%,同时保持模型通用能力。结果验证了其强鲁棒性、泛化能力与实际应用价值,为可扩展、高效的可信度修复提供了新路径。

原文摘要 · Abstract (English)

Large Language Models (LLMs) face persistent and evolving trustworthiness issues, motivating developers to seek automated and flexible repair methods that enable convenient deployment across diverse scenarios. Existing repair methods like supervised fine-tuning (SFT) and reinforcement learning with human feedback (RLHF) are costly and slow, while prompt engineering lacks robustness and scalability. Representation engineering, which steers model behavior by injecting targeted concept vectors during inference, offers a lightweight, training-free alternative. However, current approaches depend on manually crafted samples and fixed steering strategies, limiting automation and adaptability. To overcome these challenges, we propose MASteer, the first end-to-end framework for trustworthiness repair in LLMs based on representation engineering. MASteer integrates two core components: AutoTester, a multi-agent system that generates diverse, high-quality steer samples tailored to developer needs; and AutoRepairer, which constructs adaptive steering strategies with anchor vectors for automated, context-aware strategy selection during inference. Experiments on standard and customized trustworthiness tasks show MASteer consistently outperforms baselines, improving metrics by 15.36% on LLaMA-3.1-8B-Chat and 4.21% on Qwen-3-8B-Chat, while maintaining general model capabilities. MASteer demonstrates strong robustness, generalization, and practical value for scalable, efficient trustworthiness repair.

大模型可信度修复多智能体表示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。