用智能体迭代优化AI解释,让农业建议更易懂可信。
Agentic Explainable Artificial Intelligence (Agentic XAI) Approach To Explore Better Explanation
- 让大模型当智能体,反复打磨解释内容
- 改进后解释质量提升30%-33%,第3-4轮最佳
- 过犹不及,过度优化会变啰嗦抽象
可解释人工智能(XAI)能揭示数据与结果间的关系,但向非专业人士传达解释仍困难,影响对AI预测的信任。大语言模型(LLMs)有望将技术解释转化为通俗叙述,但将自主智能体(agentic AI)与XAI结合尚未探索。本研究提出一种融合SHAP解释与多模态LLM迭代优化的智能体XAI框架,用于日本26个稻田的水稻产量推荐系统。框架在11轮(第0-10轮)中逐步优化解释,由12名作物科学家和14个LLM按七项指标评估:具体性、清晰度、简洁性、实用性、情境相关性、成本考量、作物科学可信度。两组评价者均确认,解释质量平均提升30%-33%(第0轮起),峰值出现在第3-4轮。但过度迭代导致质量显著下降,揭示偏差-方差权衡:早期解释深度不足(偏差),过度迭代则引入冗余与脱离实际的抽象(方差)。研究建议采用策略性早停(正则化)以优化实用价值,挑战了‘持续改进’假设,为智能体XAI系统设计提供实证依据。
原文摘要 · Abstract (English)
Explainable artificial intelligence (XAI) enables data-driven understanding of factor associations with response variables, yet communicating XAI outputs to laypersons remains challenging, hindering trust in AI-based predictions. Large language models (LLMs) have emerged as promising tools for translating technical explanations into accessible narratives, yet the integration of agentic AI, where LLMs operate as autonomous agents through iterative refinement, with XAI remains unexplored. This study proposes an agentic XAI framework combining SHAP-based explainability with multimodal LLM-driven iterative refinement to generate progressively enhanced explanations. As a use case, we tested this framework as an agricultural recommendation system using rice yield data from 26 fields in Japan. The Agentic XAI initially provided a SHAP result and explored how to improve the explanation through additional analysis iteratively across 11 refinement rounds (Rounds 0-10). Explanations were evaluated by human experts (crop scientists) (n=12) and LLMs (n=14) against seven metrics: Specificity, Clarity, Conciseness, Practicality, Contextual Relevance, Cost Consideration, and Crop Science Credibility. Both evaluator groups confirmed that the framework successfully enhanced recommendation quality with an average score increase of 30-33% from Round 0, peaking at Rounds 3-4. However, excessive refinement showed a substantial drop in recommendation quality, indicating a bias-variance trade-off where early rounds lacked explanation depth (bias) while excessive iteration introduced verbosity and ungrounded abstraction (variance), as revealed by metric-specific analysis. These findings suggest that strategic early stopping (regularization) is needed for optimizing practical utility, challenging assumptions about monotonic improvement and providing evidence-based design principles for agentic XAI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。