专家在边界处易犯错,因表面相似掩盖本质差异。
Transitive Expert Error and Routing Problems in Complex AI Systems
- 用结构相似性与权威惯性解释专家误判机制。
- 发现跨领域路由错误与覆盖缺失导致自信的错误输出。
- 提出多专家分歧检测等可落地的改进方案。
领域专长在边界内提升判断力,却在边界处引发系统性漏洞,称为传递性专家错误(TEE),区别于达克效应。当结构相似性掩盖因果差异时,专家过度依赖表面特征(如共用词汇、模式、形式结构),忽略深层架构差异;同时,权威惯性通过社会强化与元认知失败维持信心,即使能力已超边界。该现象在共享词汇掩盖异质过程、存在即时判断压力、反馈延迟三种条件下加剧。此机制同样存在于AI路由系统(MoE、多模型编排、工具使用代理、RAG),导致路由错误(选错专家)与覆盖错误(无合适专家)。二者均产生幻觉型输出:自信、连贯、结构合理但因果错误。人类系统中机制为黑箱,而AI架构使其显式可干预。提出三类干预:路由器级多专家激活与分歧检测、专家级边界校准、训练级覆盖缺口检测。TEE具有可识别特征(路由模式、置信度-准确率脱节、领域不当内容),可用于监控与缓解。人类认知中难以解决的问题,可通过架构设计应对。
原文摘要 · Abstract (English)
Domain expertise enhances judgment within boundaries but creates systematic vulnerabilities specifically at borders. We term this Transitive Expert Error (TEE), distinct from Dunning-Kruger effects, requiring calibrated expertise as precondition. Mechanisms enabling reliable within-domain judgment become liabilities when structural similarity masks causal divergence. Two core mechanisms operate: structural similarity bias causes experts to overweight surface features (shared vocabulary, patterns, formal structure) while missing causal architecture differences; authority persistence maintains confidence across competence boundaries through social reinforcement and metacognitive failures (experts experience no subjective uncertainty as pattern recognition operates smoothly on familiar-seeming inputs.) These mechanism intensify under three conditions: shared vocabulary masking divergent processes, social pressure for immediate judgment, and delayed feedback. These findings extend to AI routing architectures (MoE systems, multi-model orchestration, tool-using agents, RAG systems) exhibiting routing-induced failures (wrong specialist selected) and coverage-induced failures (no appropriate specialist exists). Both produce a hallucination phenotype: confident, coherent, structurally plausible but causally incorrect outputs at domain boundaries. In human systems where mechanisms are cognitive black boxes; AI architectures make them explicit and addressable. We propose interventions: multi-expert activation with disagreement detection (router level), boundary-aware calibration (specialist level), and coverage gap detection (training level). TEE has detectable signatures (routing patterns, confidence-accuracy dissociations, domain-inappropriate content) enabling monitoring and mitigation. What remains intractable in human cognition becomes addressable through architectural design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。