通过分层结构解析复杂数据中的关键创新,提升高风险场景下的精准发现能力。
HH-SAE: Discovering and Steering Hierarchical Knowledge of Complex Manifolds

- 构建三级分层架构,将数据从背景中分离出生理与行为模式。
- 在反欺诈任务中实现0.9156的跨域零样本AUC,显著优于现有方法。
- 适合医疗、金融等对精度要求高的复杂系统分析场景。
高维、高风险领域中的罕见语义创新常被密集的背景信息掩盖,我们称之为‘特征密度冲突’。为此提出混合分层自编码器(HH-SAE),将流形分解为上下文(L₀)、原子(f₁)和组合(f₂)三层结构。在多种不同流形上评估显示,HH-SAE能有效‘解构’行政临床标签,还原为生理模式,并在反欺诈任务中达到0.9156的跨域零样本AUC。路径消融实验表明,移除上下文减法会导致性能下降13.46%。知识引导生成实现了比当前最优生成器高出9.9%的AUPRC,证明了该模型能优先捕捉高层次机制创新而非环境代理,从而在高风险环境中实现高精度发现。
原文摘要 · Abstract (English)
Rare semantic innovations in high-dimensional, mission-critical domains are often obscured by dense background contexts, a challenge we define as \textit{feature density conflict}. We introduce the \textbf{Hybrid Hierarchical SAE (HH-SAE)} to resolve this by factorizing manifolds into a nested hierarchy of \textbf{Contextual} ($L_0$), \textbf{Atomic} ($f_1$), and \textbf{Compository} ($f_2$) tiers. Evaluating across disparate manifolds, HH-SAE demonstrates superior resolution by \textbf{``fracturing'' administrative clinical labels into physiological modes} and achieving a peak \textbf{cross-domain zero-shot AUC of 0.9156 in fraud detection}. Path ablation confirms the architecture's structural necessity, revealing a 13.46\% utility collapse when contextual subtraction is removed. Finally, knowledge-steered synthesis achieves a +9.9\% AUPRC lift over state-of-the-art generators, proving that HH-SAE effectively prioritizes high-order mechanistic innovation over environmental proxies to enable high-precision discovery in high-stakes environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。