arXiv:2601.05637cs.AIcs.LG2026-01被引 1

提出一套理论工具,判断生成模型能否被精确控制。

GenCtrl -- A Formal Controllability Toolkit for Generative Models

  • 用对话场景建模人机交互,定量估算模型可控区域。
  • 在不假设分布的前提下,给出误差的严格数学保证。
  • 发现控制效果极易受设置影响,需先评估控制极限。

随着生成模型广泛应用,对生成过程进行细粒度控制的需求日益迫切。尽管从提示到微调的控制方法层出不穷,但一个根本问题仍未解决:这些模型是否真的可控?本文提出一个理论框架,以人机交互为控制过程,设计新算法估计对话场景下模型的可控集合。我们首次提供无需分布假设、仅依赖输出有界的误差概率保证,适用于任意黑箱非线性控制系统(即任意生成模型)。在语言模型与文本到图像生成任务中,实证结果表明模型可控性出人意料地脆弱,且高度依赖实验设置。这凸显了严谨可控性分析的必要性,推动研究重心从盲目尝试控制转向理解其基本边界。

原文摘要 · Abstract (English)

As generative models become ubiquitous, there is a critical need for fine-grained control over the generation process. Yet, while controlled generation methods from prompting to fine-tuning proliferate, a fundamental question remains unanswered: are these models truly controllable in the first place? In this work, we provide a theoretical framework to formally answer this question. Framing human-model interaction as a control process, we propose a novel algorithm to estimate the controllable sets of models in a dialogue setting. Notably, we provide formal guarantees on the estimation error as a function of sample complexity: we derive probably-approximately correct bounds for controllable set estimates that are distribution-free, employ no assumptions except for output boundedness, and work for any black-box nonlinear control system (i.e., any generative model). We empirically demonstrate the theoretical framework on different tasks in controlling dialogue processes, for both language models and text-to-image generation. Our results show that model controllability is surprisingly fragile and highly dependent on the experimental setting. This highlights the need for rigorous controllability analysis, shifting the focus from simply attempting control to first understanding its fundamental limits.

可控性分析生成模型理论框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。