用经济学思维提升视觉语言模型的安全与效率。
EcoAlign: An Economically Rational Framework for Efficient LVLM Alignment
- 将对齐视为理性决策过程,动态权衡安全、效用和成本。
- 在6个数据集上以更低计算开销达到顶尖安全与性能。
- 适合关注模型安全性与推理效率的开发者和研究者。
大型视觉语言模型(LVLMs)具备强大推理能力,但存在复杂的越狱漏洞。对齐不仅关乎安全,更涉及经济效率。现有方法难以平衡安全性、实用性与运行成本。关键问题在于仅关注最终输出(过程盲视),导致大量计算资源被浪费在不安全的推理过程中,使有害推理可通过表面合理的解释规避简单的安全评分机制。为此,我们提出EcoAlign,一个推理时框架,将对齐重构为经济理性的搜索过程,将LVLM视为有限理性的代理。EcoAlign逐步扩展思维图谱,并使用前瞻函数(类比净现值)动态评估动作的安全性、效用与成本,结合剩余预算进行评分。为防止欺骗,采用最弱环节原则强制路径安全。在3个闭源和2个开源模型、6个数据集上的实验表明,EcoAlign在更低计算成本下实现或超越当前最优的安全性与实用性,提供了一条原则性强、经济高效的鲁棒对齐路径。
原文摘要 · Abstract (English)
Large Vision-Language Models (LVLMs) exhibit powerful reasoning capabilities but suffer sophisticated jailbreak vulnerabilities. Fundamentally, aligning LVLMs is not just a safety challenge but a problem of economic efficiency. Current alignment methods struggle with the trade-off between safety, utility, and operational costs. Critically, a focus solely on final outputs (process-blindness) wastes significant computational budget on unsafe deliberation. This flaw allows harmful reasoning to be disguised with benign justifications, thereby circumventing simple additive safety scores. To address this, we propose EcoAlign, an inference-time framework that reframes alignment as an economically rational search by treating the LVLM as a boundedly rational agent. EcoAlign incrementally expands a thought graph and scores actions using a forward-looking function (analogous to net present value) that dynamically weighs expected safety, utility, and cost against the remaining budget. To prevent deception, path safety is enforced via the weakest-link principle. Extensive experiments across 3 closed-source and 2 open-source models on 6 datasets show that EcoAlign matches or surpasses state-of-the-art safety and utility at a lower computational cost, thereby offering a principled, economical pathway to robust LVLM alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。