用拓扑方法分析大模型对模糊问题的内部状态,提升识别与回应质量。
The Topology of Ill-Posed Questions: Persistent Homology for Detection and Steering in LLMs

- 将提示词的隐藏状态建模为点云,用零维持久同调分析其几何结构。
- 在三个模型上提升模糊问题分类准确率,最高达88.5%,并改善可接受回答率。
- 通过拓扑相似性引导响应修正,适合需要严谨推理或澄清的场景。
模糊问题(如歧义、信息不足或矛盾)常导致大语言模型无有效答案或多解,构成挑战。现有方法多基于输出分析且聚焦特定子类。本文探究不同类型的模糊性是否可在统一的模型内部状态拓扑中表征,并用于引导响应行为。将每个变换器层中提示词的上下文隐藏状态视为点云,利用有限零维持久同调刻画其几何特征。每层用三个紧凑描述符总结:平均有限寿命、归一化寿命熵、最大寿命集中度。跨层拼接后形成问题的拓扑表示。进一步提出拓扑条件激活引导,检索拓扑相似示例并构建针对具体查询的激活干预,促进来源意识的澄清或放弃回答。在三个开源大模型上,拓扑特征在模糊性分类任务中持续优于提示基和池化隐藏状态基线,在AmbigQA上准确率从67.4%提升至78.9%,SituatedQA上从79.9%升至88.5%,CLAMBER 9分类任务从57.6%增至69.6%。拓扑引导使平均可接受响应率从61.4%升至70.6%,可落地的可接受响应率从11.9%增至16.4%。结果表明,持久同调既提供可解释的模糊性表征,也构成有效的定向响应引导机制。
原文摘要 · Abstract (English)
Ill-posed questions, including ambiguous, underspecified, or contradictory queries, may admit no valid answer or multiple plausible answers, posing a challenge for large language models (LLMs). Existing approaches largely analyze ill-posedness through model outputs and often focus on specific subclasses. We investigate whether diverse sources of ill-posedness can be represented within a unified topology of LLM internal states and whether this structure can be used to steer response behavior. We model the contextual hidden states of prompt tokens at each transformer layer as a point cloud and characterize its geometry using finite zero-dimensional persistent homology. Each layer is summarized by three compact descriptors: mean finite lifetime, normalized lifetime entropy, and largest-lifetime concentration. Concatenating these descriptors across layers yields a topology representation of the question. We further introduce topology-conditioned activation steering, which retrieves topologically similar examples and constructs query-specific activation interventions that encourage source-aware clarification or abstention. Across three open-weight LLMs, topology features consistently outperform prompt-based and pooled-hidden-state baselines for ill-posedness classification, improving average accuracy from \(67.4\%\) to \(78.9\%\) on AmbigQA, from \(79.9\%\) to \(88.5\%\) on SituatedQA, and from \(57.6\%\) to \(69.6\%\) on CLAMBER 9-way classification. Topology-conditioned steering increases the average total acceptable response rate from \(61.4\%\) to \(70.6\%\) and grounded acceptable responses from \(11.9\%\) to \(16.4\%\). These results show that persistent homology provides both an interpretable representation of ill-posedness and an effective mechanism for targeted response steering.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。