arXiv:2512.21699cs.AI2025-12被引 12

通过多模型共识机制,让智能体决策更可靠可解释。

Towards Responsible and Explainable AI Agents with Consensus-Driven Reasoning

  • 多个模型独立推理后由专门代理整合意见,增强判断可靠性。
  • 实测显示该架构显著提升任务鲁棒性与透明度,减少幻觉和偏见。
  • 适合需要安全可控的生产级智能体系统,如医疗、金融决策场景。

智能体人工智能通过协调大语言模型(LLM)、视觉语言模型(VLM)、工具与外部服务,实现多步骤任务的自主规划与执行。尽管能力强大,但自主性增强带来了可解释性、责任归属、鲁棒性与治理等关键挑战,尤其在影响实际决策时。现有系统多关注功能与扩展性,缺乏对决策逻辑的透明支持与责任约束机制。本文提出基于多模型共识与推理层治理的负责任(RAI)与可解释(XAI)智能体架构,通过异构的LLM与VLM代理在共享输入下独立生成候选输出,显式暴露不确定性、分歧与不同解读。一个专用推理代理对这些输出进行结构化整合,强制执行安全与政策约束,缓解幻觉与偏见,并生成可审计、有证据支持的决策。可解释性通过跨模型对比与中间输出保留实现,责任则通过中心化推理层控制与代理级约束保障。我们在多个真实世界智能体工作流中评估该架构,证明共识驱动推理能有效提升鲁棒性、透明度与运营信任,适用于需要自主、可扩展又负责任的生产级应用。

原文摘要 · Abstract (English)

Agentic AI represents a major shift in how autonomous systems reason, plan, and execute multi-step tasks through the coordination of Large Language Models (LLMs), Vision Language Models (VLMs), tools, and external services. While these systems enable powerful new capabilities, increasing autonomy introduces critical challenges related to explainability, accountability, robustness, and governance, especially when agent outputs influence downstream actions or decisions. Existing agentic AI implementations often emphasize functionality and scalability, yet provide limited mechanisms for understanding decision rationale or enforcing responsibility across agent interactions. This paper presents a Responsible(RAI) and Explainable(XAI) AI Agent Architecture for production-grade agentic workflows based on multi-model consensus and reasoning-layer governance. In the proposed design, a consortium of heterogeneous LLM and VLM agents independently generates candidate outputs from a shared input context, explicitly exposing uncertainty, disagreement, and alternative interpretations. A dedicated reasoning agent then performs structured consolidation across these outputs, enforcing safety and policy constraints, mitigating hallucinations and bias, and producing auditable, evidence-backed decisions. Explainability is achieved through explicit cross-model comparison and preserved intermediate outputs, while responsibility is enforced through centralized reasoning-layer control and agent-level constraints. We evaluate the architecture across multiple real-world agentic AI workflows, demonstrating that consensus-driven reasoning improves robustness, transparency, and operational trust across diverse application domains. This work provides practical guidance for designing agentic AI systems that are autonomous and scalable, yet responsible and explainable by construction.

智能体可解释责任多模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。