arXiv:2607.29008stat.MLcs.LG2026-07被引 1

用拓扑方法测试模型语义对齐,让黑箱模型更可解释。

Persistent Convolution: A Topological Framework for AI Alignment Testing and Semantic Space Characterization

论文配图:Persistent Convolution: A Topological Framework for AI Alignment Testing and Semantic Space Characterization
图 1 · 摘自论文原文
  • 基于拓扑学构建多模态对齐检测框架
  • 能有效刻画语义结构与概念分离程度
  • 适合模型选型与部署时的可解释性评估

现代复杂AI模型注重性能而牺牲可解释性,导致测试困难。本研究提出一种基于拓扑的多模态对齐测试方法,通过在模型嵌入空间上进行形式化统计检验,实现对语义结构、概念分离和知识图谱对齐的稳健表征。该方法利用人工标注的知识结构进行模型对比,提升模型选择与部署的可解释性。针对输入空间规模大带来的挑战,该框架提供了一种超越传统输出分析的对齐验证手段,并与可能性理论及统一决策理论框架自然衔接,为从数据到部署的全链条提供可解释支持。

原文摘要 · Abstract (English)

Modern opaque AI models prize performance over interpretability, which makes testing difficult. However, formal statistical tests conducted on a model's embedding space can provide robust characterizations of semantic structure, concept separation, and knowledge graph alignment. Model developers would benefit from a model comparison technique that leverages human-curated knowledge structures to test alignment. The scale of the input space for even relatively simple tasks motivates the need for alignment checks that augment standard outcome reasoning. This work develops and demonstrates a topology-based multi-modal alignment test to make deployment, selection, and comparison of opaque models more interpretable. These methods also offer an intuitive connection to possibility theory and a unified decision theoretic framework from data to deployment.

可解释性拓扑方法模型测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。