arXiv:2602.00638cs.CL2026-02

让语言模型的内部表示可解释、可控制,实现语义精准操作。

Formal Semantic Control over Language Models

  • 在变分自编码器框架下重塑隐空间几何,实现语义特征解耦与操控。
  • 在句子生成和推理任务中均实现对特定语义的局部可控生成。
  • 适合研究模型可解释性、可控生成及逻辑推理的学者参考。

本论文致力于提升语言表示的学习能力,使语言模型的隐空间具备更强的语义与几何可解释性,并通过精心设计的隐空间结构实现局部化、准符号化的组合式控制。研究基于变分自编码器框架,探索两个互补方向:(i) 句子级学习与控制:在隐空间中解耦并操纵特定语义特征,引导句子生成,以解释性文本为测试场景;(ii) 推理级学习与控制:在隐空间中隔离并引导推理行为,以控制自然语言推理(NLI)。在此方向上,聚焦于解释性NLI任务,即通过两个前提(解释)推导结论。整体目标是推动语言模型的内部语义表示走向系统可解释、精确构造与可靠引导。本文提出一系列新理论框架与实用方法,并通过实验验证,显著提升了自然语言隐空间在语义可解释性与可控性方面的表现。

原文摘要 · Abstract (English)

This thesis advances semantic representation learning to render language representations or models more semantically and geometrically interpretable, and to enable localised, quasi-symbolic, compositional control through deliberate shaping of their latent space geometry. We pursue this goal within a VAE framework, exploring two complementary research directions: (i) Sentence-level learning and control: disentangling and manipulating specific semantic features in the latent space to guide sentence generation, with explanatory text serving as the testbed; and (ii) Reasoning-level learning and control: isolating and steering inference behaviours in the latent space to control NLI. In this direction, we focus on Explanatory NLI tasks, in which two premises (explanations) are provided to infer a conclusion. The overarching objective is to move toward language models whose internal semantic representations can be systematically interpreted, precisely structured, and reliably directed. We introduce a set of novel theoretical frameworks and practical methodologies, together with corresponding experiments, to demonstrate that our approaches enhance both the interpretability and controllability of latent spaces for natural language across the thesis.

语义控制隐空间可解释性生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。