用进化稳定机制防范AI代理欺诈,提升去中心化服务可信度。
Ev-Trust: An Evolutionarily Stable Trust Mechanism for Decentralized LLM-Based Multi-Agent Service Economies
- 通过语义验证与信任信号嵌入,实现精准欺诈识别
- 实验显示恶意参与率降60%,欺诈服务率降50%
- 适合构建安全可靠的AI代理协作系统的研究者
去中心化的基于大模型的多智能体服务经济面临三大漏洞:欺诈成本降低、服务质量难评估、服务内容不稳定。这些漏洞叠加可能导致群体性信任崩溃和短视策略泛滥。我们提出Ev-Trust,一种进化稳定的信任机制,通过三项设计应对:利用请求方语义理解的交叉验证门来评估回复有效性;基于方差标准化的漂移度量,分离内生随机性与真实行为异常;将信任信号嵌入预期收益函数,使可信成为进化生存优势。基于带有噪声最优响应微观基础的复制者动态,我们证明了合作演化稳定策略的渐近稳定性,并推导出维持合作均衡的显式阈值条件。我们在包含至少100个异构大模型驱动智能体、覆盖七种行为类型的100轮模拟中评估了Ev-Trust,使用TruthfulQA和TriviaQA两个事实问答基准。相比基于传递性信任聚合、强化学习声誉和纯进化模仿的基线方法,Ev-Trust将恶意智能体参与率降低约60%,欺诈服务率降低约50%,并在30%对抗性突变下保持稳定的信任分化。结果表明,结合语义信任评估与进化激励可为去中心化大模型多智能体系统提供稳固的合作保障。
原文摘要 · Abstract (English)
Decentralized LLM-based multi-agent service economies face three vulnerabilities that undermine traditional trust mechanisms: reduced cost of fraud, difficulty in evaluating service quality, and instability of service content. These compounding vulnerabilities can trigger population-level trust collapse and the proliferation of short-sighted strategies. We propose Ev-Trust, an evolutionarily stable trust mechanism that addresses these vulnerabilities through three targeted designs: a cross-validation gate leveraging requestor semantic comprehension to assess response validity, a variance-standardized drift measure filtering endogenous stochasticity from genuine behavioral anomalies, and an embedding of trust signals into the expected revenue function that converts trustworthiness into an evolutionary survival advantage. Based on replicator dynamics with a noisy best response micro-foundation, we prove the asymptotic stability of cooperative evolutionarily stable strategies and derive explicit threshold conditions for maintaining cooperative equilibria. We evaluate Ev-Trust through 100-round simulations with at least 100 heterogeneous LLM-driven agents covering seven behavioral types. The experiments are conducted on TruthfulQA and TriviaQA, two factual question-answering benchmarks. Compared to baselines based on transitive trust aggregation, reinforcement-learning reputation, and pure evolutionary imitation, Ev-Trust reduces malicious agent participation by approximately 60%, suppresses the fraudulent service rate by approximately 50%, and maintains stable trust differentiation under a 30% adversarial mutation. These results demonstrate that coupling semantic trust evaluation with evolutionary incentives provides a principled foundation for securing cooperation in decentralized LLM-based multi-agent systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。