arXiv:2510.03271cs.LGcs.AI2025-10

提出新方法精准逼近大模型决策边界,仅需少量样本即可实现高精度分析。

Decision Potential Surface: A Theoretical and Practical Approximation of Large Language Model Decision Boundary

  • 引入决策势面(DPS)概念,从分类置信度捕捉决策边界潜力。
  • 首次提出K-DPS算法,仅用K个序列样本就可近似决策边界,误差极小。
  • 理论证明误差可调控,适合研究大模型行为与可解释性的人群使用。

决策边界是机器学习模型中对两类分类概率相等的输入子空间,对揭示模型本质特性与行为解释至关重要。尽管大语言模型(LLMs)的决策边界分析近年备受关注,但因序列输出空间庞大且具有自回归特性,主流LLMs的决策边界构造仍面临计算不可行的问题。本文提出决策势面(DPS),一种用于分析LLM决策特性的新概念。DPS基于每个输入区分不同类别的置信度,自然刻画了决策边界的潜在特性。我们证明DPS的零高度等高线等价于LLM的决策边界,其包围区域代表决策区域。借助DPS,我们首次在文献中提出实用的决策边界近似算法K-DPS,仅需K个有限序列样本,即可以可忽略的误差近似LLM的决策边界。我们理论上推导了K-DPS与理想DPS之间绝对误差、期望误差及误差集中度的上界,表明这些误差可随采样次数进行权衡。

原文摘要 · Abstract (English)

Decision boundary, the subspace of inputs where a machine learning model assigns equal classification probabilities to two classes, is pivotal in revealing core model properties and interpreting behaviors. While analyzing the decision boundary of large language models (LLMs) has attracted increasing attention recently, constructing it for mainstream LLMs remains computationally infeasible due to the enormous sequence-level output spaces and the autoregressive nature of LLMs. To address this issue, in this paper we propose Decision Potential Surface (DPS), a new notion for analyzing the properties of LLM decisions. DPS is derived from the confidence in distinguishing different classes for each input, which naturally captures the potential of the decision boundary. We prove that the zero-height isohypse in DPS is equivalent to the decision boundary of an LLM, with enclosed regions representing decision regions. By leveraging DPS, for the first time in the literature, we propose a practical decision boundary approximation algorithm, namely K-DPS, which only requires only K finite sequence samples to approximate an LLM's decision boundary with negligible error. We theoretically derive the upper bounds for the absolute error, expected error, and the error concentration between K-DPS and the ideal DPS, demonstrating that such errors can be traded off against sampling times.

大模型决策边界可解释性理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。