用本体增强蒸馏打造企业级语言模型,验证其在金融领域的准确性与可审计性。
Ontology-Amplified Distillation and Contextuality Auditing for Sovereign Enterprise Language Models: A Combined Proof-of-Mechanism and Negative-Results Method Study
- 通过本体引导的微调与偏好优化,让小模型在本地单卡上复现大模型能力。
- 在40个越南金融任务中,90%的任务实现准确本体对齐,覆盖率达95%。
- 提出可审计的上下文一致性检测方法,指导模型决策是否需人工介入。
受数据驻留法规约束的金融机构需要可由租户自主拥有、运行于机构内部的语音模型。本文结合两项相关研究,提出机制与控制并重的方法论。首先,开展低算力验证实验:基于47对合成英文跨领域偏好数据,在单台Apple M5 Max上对Qwen3.6-27B学生模型进行本体增强蒸馏,通过监督微调与本体引导的直接偏好优化(DPO)训练。在40个保留的越南金融任务中,该模型成功对齐36项任务(对齐率0.90),平均本体术语覆盖率达0.95(下限为0.50),与GPT-5基准表现相当。但统计效力不足,无法确认等效性——配对差异95%置信区间覆盖±4任务,且未验证预注册的“学生应超越前沿”预测。其次,整合一种上下文一致性审计方法用于企业代理路由。在独立的负结果试点中,本地Qwen运行及显式标注的Gemma复现检查均显示,所有1.3阶段组的上下文一致性度量(Contextuality-by-Default)均为零;有效信号来自直接影响力与构念耦合,而非残留上下文一致性。两研究共同构建了本体驱动建模机制与治理诊断框架,用于判断模型分歧是否需触发提示标准化、多代理合成或人工审查。证据不支持部署可行性、安全性、优越性、统计等效性或正向上下文一致性的路由规则。
原文摘要 · Abstract (English)
Regulated financial institutions operating under data-residency rules need tenant-owned language models that can run inside the institution's perimeter. This paper combines two related FAOS studies into one mechanism-and-control article. First, it reports a reduced-power proof-of-mechanism study of ontology-amplified distillation: a Qwen3.6-27B student is adapted to the Foundation AgenticOS ontology through supervised fine-tuning on frontier-teacher trajectories and ontology-grounded direct preference optimization (DPO), trained locally on a single Apple M5 Max from 47 synthetic, English-language, cross-domain preference pairs. On 40 held-out Vietnamese financial-domain tasks, the distilled student grounds 36 of 40 tasks (grounded rate 0.90; mean ontology term-coverage r_onto = 0.95 on a metric floored at 0.50), equal to the GPT-5 frontier baseline, which also grounds 36 of 40. The outcome is underpowered to establish equivalence: the paired-difference 95% confidence interval spans +/-4 tasks, and the run does not test or show the pre-registered amplification prediction that the student should exceed the frontier. Second, the paper consolidates a contextuality-audit method for enterprise-agent routing. In a separate negative-results pilot, the corrected canonical Contextuality-by-Default degree is zero for all Phase 1.3 groups in both the local-Qwen run and an explicitly labeled Gemma replication check; the useful signal is direct influence and construct coupling, not surviving residual contextuality. Together, the studies pair an ontology-grounded model-building mechanism with a governance diagnostic for deciding when apparent disagreement should trigger prompt standardization, multi-agent synthesis, or human review. The evidence supports neither deployability, safety, superiority, statistical equivalence, nor a contextuality-positive routing rule.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。