用实时测试取代文字描述,让多智能体系统更安全地分配任务。
Linguistic Firewall: Geometry as Defense in Multi-Agent Systems Routing

- 通过动态测试智能体真实能力,生成非文本行为算子进行路由。
- 对描述注入攻击的误判率接近零,比基线低67.3%以上。
- 适合需要高安全性的智能体协作场景,如金融、医疗等敏感领域。
大型语言模型的快速集成推动了多智能体系统(MAS)的发展,各专业智能体协同完成复杂任务。现有路由机制依赖未经验证的间接代理(如文本自我描述或静态表征)评估智能体能力,导致其预期画像与实际表现脱节,带来严重安全风险。恶意智能体可伪装能力或隐藏后门,逃避常规分析和静态学习技术检测。本文提出ANTAP(自动非文本智能体选择器),一种以评估驱动的路由架构,摒弃间接代理,改用动态查询获取真实能力。通过将性能提炼为共享语义空间中的固定行为算子,在推理时仅通过非文本代数投影完成路由,构建‘语言防火墙’,使基于元数据的攻击无法表达。实验表明,针对基于描述的注入攻击,ANTAP的误判率接近零,而基线达67.3%以上;面对自适应嵌入攻击,其误判率比嵌入基线降低20%,且天然抵御描述操纵。
原文摘要 · Abstract (English)
The rapid integration of Large Language Models (LLMs) has driven the evolution of Multi-Agent Systems (MAS), where specialized agents collaborate to execute complex workflows. Effective orchestration in these environments requires robust routing mechanisms to efficiently allocate tasks to the most suitable agent. However, existing routers fundamentally rely on unverified proxies, ranging from textual self-descriptions to static surrogate representations, to gauge an agent's competence. This reliance on non-empirical data creates a critical gap between an agent's projected profile and its actual operational capabilities, introducing severe security vulnerabilities. Malicious agents can easily misrepresent their proficiencies or harbor covert backdoors that evade both standard external analysis and static representation-learning techniques. In this work, we introduce ANTAP (Automatic Non-Textual Agent Picker), an evaluation-driven routing architecture that discards indirect proxies in favor of active capability testing. By dynamically querying agents to ascertain their true competencies empirically, ANTAP distills performance into fixed behavioral operators within a shared semantic space. At inference time, routing is performed via a purely non-textual algebraic projection, establishing a "linguistic firewall" that renders metadata-based attacks inexpressible. In our experiments, ANTAP achieves near-zero ASR against description-based injection attacks, compared to 67.3\% and above for the description-based router baseline. Against adaptive embedding attacks, ANTAP achieves substantially lower ASR than the embedding-based baseline, with a 20\% reduction, while remaining resilient to description manipulation by design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。