arXiv:2511.03121cs.CLcs.AI2025-11被引 11

用控制屏障函数让大模型生成符合用户期望的文本

Control Barrier Function for Aligning Large Language Models

  • 通过控制屏障函数实时干预生成文本,不需微调原模型
  • 在多个数据集上实现90%以上的正面文本生成率
  • 适合需要快速对齐大模型输出的场景,如安全对话系统

本文提出一种基于控制的框架,利用控制屏障函数(CBF)确保大语言模型(LLM)生成用户期望的文本。该框架将CBF安全过滤器应用于基线LLM生成的预测词元,进行实时干预。该安全过滤器具有两大优势:一是作为可插拔组件,无需微调基线模型即可用于对齐;二是若存在评估模型衡量期望对齐效果,可直接用于滤波器设计。整个文本生成系统基于开源语言模型实现,旨在生成积极正面的文本。

原文摘要 · Abstract (English)

This paper proposes a control-based framework for aligning large language models (LLMs) by leveraging a control barrier function (CBF) to ensure user-desirable text generation. The presented framework applies the CBF safety filter to the predicted token generated from the baseline LLM, to intervene in the generated text. The safety filter includes two significant advantages: this safety filter is an add-on type, allowing it to be used for alignment purposes without fine-tuning the baseline LLM, and if there is an evaluation model regarding the desired alignment, it can be directly applied to the filter design. The overall text-generation system is implemented with open-source language models, aiming to generate positive text.

大模型对齐控制屏障文本生成安全过滤

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。