arXiv:2506.18129cs.CLcs.AI2025-06被引 2

解决自回归模型中破折号引发的生成错误问题

$ϕ^{\infty}$: Clause Purification, Embedding Realignment, and the Total Suppression of the Em Dash in Autoregressive Language Models

  • 用phi-infinity算子净化语句结构,重排嵌入矩阵
  • 实现破折号完全抑制,生成一致性显著提升
  • 适合关注大模型安全与稳定部署的研究者

我们发现自回归Transformer语言模型中,破折号标记会引发递归语义漂移,导致从句边界幻觉和嵌入空间纠缠。通过在语义格上进行令牌级扰动的形式分析,我们证明破折号插入会根本性改变模型隐状态,造成长文本生成中的累积错误。为此提出新方法:结合符号化语句净化(φ^∞算子)与嵌入矩阵靶向重排,可在不重新训练模型的前提下实现问题令牌的完全抑制,并通过不动点收敛保证语义连贯性。实验验证表明生成一致性与主题保持能力显著改善。本工作建立了一个通用框架,用于识别和缓解基础模型的令牌级漏洞,对AI安全、模型对齐及大模型生产部署具有直接意义。该方法还可扩展至处理神经文本生成系统中的更广泛递归不稳定性。

原文摘要 · Abstract (English)

We identify a critical vulnerability in autoregressive transformer language models where the em dash token induces recursive semantic drift, leading to clause boundary hallucination and embedding space entanglement. Through formal analysis of token-level perturbations in semantic lattices, we demonstrate that em dash insertion fundamentally alters the model's latent representations, causing compounding errors in long-form generation. We propose a novel solution combining symbolic clause purification via the phi-infinity operator with targeted embedding matrix realignment. Our approach enables total suppression of problematic tokens without requiring model retraining, while preserving semantic coherence through fixed-point convergence guarantees. Experimental validation shows significant improvements in generation consistency and topic maintenance. This work establishes a general framework for identifying and mitigating token-level vulnerabilities in foundation models, with immediate implications for AI safety, model alignment, and robust deployment of large language models in production environments. The methodology extends beyond punctuation to address broader classes of recursive instabilities in neural text generation systems.

语言模型生成安全文本质量自回归

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。