arXiv:2605.08964cs.LG2026-05

提升机器学习系统可靠性与可问责性,覆盖模型到智能体全链条

Trustworthy AI: Ensuring Reliability and Accountability from Models to Agents

论文配图:Trustworthy AI: Ensuring Reliability and Accountability from Models to Agents
图 1 · 摘自论文原文
  • 用核方法实现复杂子群体的多精度预测,减少偏见与任意决策
  • 提出水印机制,在文本生成中平衡检测率与内容质量,最优水印策略经理论推导
  • 首个全大模型驱动供应链模拟器,验证智能体性能超越人类团队67%成本降低

本论文针对机器学习系统从预测模型向生成模型和自主智能体演进带来的可信性挑战,提出具有理论保障的算法。为缓解传统模型中的偏见与任意性,引入基于核的方法,实现对复杂子群体的多精度预测,超越传统人口统计分类。针对预测多重性问题,提出可解释的统一框架。在生成式AI中,通过水印技术确保内容可溯源,刻画了水印检测与文本失真间的信息论权衡,结合最优传输与编码理论推导出最优水印策略。实验表明该水印在语言生成与编码任务中实现更优的检测-质量折衷。最后,构建首个全大模型驱动的供应链多智能体仿真系统,评估结果显示大模型智能体性能显著优于人工团队,成本降低最高达67%,但也带来高代价尾部风险等系统性隐患。

原文摘要 · Abstract (English)

In this thesis, we develop algorithms with theoretical guarantees for ensuring reliability and accountability of Machine Learning (ML) systems. As ML systems evolve from predictive models to generative models and autonomous agents, the landscape of trustworthy AI has shifted. This thesis introduces tools grounded in information theory, optimization, and statistical learning to mitigate bias, reduce arbitrary decisions, ensure content provenance, and evaluate LLM-driven agents in autonomous settings. Towards mitigating bias and arbitrariness in traditional ML models, we introduce a kernel-based method to achieve multiaccuracy across complex subpopulations that traditional demographic categories may overlook. We also develop methods to address predictive multiplicity, where equally accurate models yield conflicting individual predictions. We ensure the accountability in generative AI through watermarking large language models (LLMs). We characterize the information-theoretic trade-off between watermark detection and text distortion and derive optimal watermarking strategies by leveraging optimal transport and coding theory. Empirical evaluations show our watermarks achieve a superior detection-quality tradeoff across language generation and coding tasks. Finally, we evaluate autonomous LLM agents in multi-agent environments through the first simulator of a fully LLM-driven supply chain. LLM agents offer significant performance gains, outperforming human teams and reducing costs by up to 67%, but also introduce systemic risks, including costly tail events.

可信AI大模型水印智能体评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。