提出可自检的几何信念接口验证框架,提升企业级系统可信度。
BeTaL-GBI: Admission-Aware Benchmark Tuning and Full-Stack Verification of Geometric Belief Interfaces
- 引入大模型驱动的基准调优,分离格式接纳与任务性能评估
- 在512个合成任务中实现100%矛盾检测与零误通过率
- 支持政策版本化与自我审计,适合高可信系统研发团队
验证框架的可信性在于能否暴露自身主张中的错误,而不仅限于模型输出。GBI-DCSE v3 揭露了架构主张的偏差:报告的 Fisher 值 epsilon ~ 0.066 仅在区间 [epsilon, 3, 4, 5] 满足 kappa^2 <= 10^4 预算,而完整区间 [epsilon, 20]^4 要求 epsilon ~ 0.326472。本研究探讨企业级验证架构是否能隔离接口故障、任务能力、策略合规性与控制完整性,并保持主张可审计。BoundaryBench v0.1 基线显示:Qwen3-4B-Instruct-2507 完成 768 次冻结执行,但 0% 通过合约(369 次解析失败,399 次验证失败),限制下游选择性指标。本研究评估三项改进:首先,BeTaL-GBI v0.2 在 2,218,750,380 个网格点上进行大模型循环的基准调优,分离格式接纳(rho_adm = N_admitted/N)与条件性能(rho_task = N_verified/N_admitted)。经模式修复后,模型无关反馈搜索实现 2.87% 的平均保留目标差距,优于非反馈基线(13.61%,11.46%)。其次,GBI v2 将静态密钥替换为无参考见证状态 W 与策略 P。在 512 个合成任务中,16 门策略检测到全部 116 个注入严重矛盾,接受全部 99 个干净记录(广义分母:4.27%)。幻觉生成器与证据伪造代理均被零沉默放行阻断。第三,GBI-DCSE v3 将 99 条主张映射为机器可读证据:95/96 可测试主张通过,执行 148 次独立检查无失败。该框架在 62 种配置下验证签名账本、PBFT 仲裁与安全区伪造。在合成条件下,GBI-DCSE 是可选择、策略版本化、自我审计的测试与路由基础架构。
原文摘要 · Abstract (English)
A verification substrate is more credible when exposing errors in its own claims, not just model outputs. GBI-DCSE v3 falsified an architectural claim: the reported Fisher value epsilon ~ 0.066 satisfies the kappa^2 <= 10^4 budget only on the slice [epsilon, 3, 4, 5], while the full box [epsilon, 20]^4 requires epsilon ~ 0.326472. This erratum highlights whether an enterprise verification architecture can isolate interface failure, task competence, policy admissibility, and control integrity while keeping claims auditable. BoundaryBench v0.1 established the baseline: Qwen3-4B-Instruct-2507 completed 768 frozen executions, but 0% cleared the contract (369 failed parsing, 399 failed validation), limiting downstream selectivity metrics. This companion study evaluates three successive improvements. First, BeTaL-GBI v0.2 applies Benchmark Tuning with an LLM-in-the-loop over 2,218,750,380 grid points, separating format admission from conditional performance (rho_adm = N_admitted/N; rho_task = N_verified/N_admitted). Following schema repair, a model-free feedback search achieves a 2.87% mean held-out target gap, outperforming non-feedback baselines (13.61%, 11.46%). Second, GBI v2 swaps static keys for a reference-independent witness state W and policy P. Across 512 synthetic tasks, a 16-gate policy detects all 116 injected severe contradictions and accepts all 99 clean records (broad denominator: 4.27%). Hallucinator and evidence-forger surrogates are blocked with zero silent promotions. Third, GBI-DCSE v3 maps 99 claims to machine-readable evidence: 95 of 96 testable claims pass, with 148 standalone checks executed without failure. The harness exercises signed ledgers, PBFT quorums, and enclave forgery across 62 configurations. Under synthetic conditions, GBI-DCSE is a selective, policy-versioned, self-auditing test and routing substrate.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。