用形式化论证解决多智能体需求协商中的冲突,提升可审计性。
ArgRE: Formal Argumentation for Conflict Resolution in Multi-Agent Requirements Negotiation
- 将需求提案建模为论据,冲突表现为攻击关系,用形式化语义求解可接受集。
- 决策解释得分显著更高(4.32 vs. 3.07),合规覆盖率达84.7%。
- 适合需要可追溯、可审查的需求协商场景,如安全关键系统。
随着软件系统复杂度提升,需平衡日益增多的相互竞争的质量属性,例如传感器融合验证的安全要求可能与紧缩的规划周期预算冲突。现有基于多智能体大语言模型的框架通过分配专用智能体实现目标平衡,但其冲突解决依赖启发式方法,需求隐式聚合且无显式接受或拒绝,限制了在受监管领域的可审计性。本文提出ArgRE,一个嵌入Dung风格抽象论证的多智能体需求协商系统。每个提案、批判与修正均被建模为论据,冲突表示为有向攻击关系,可接受论据集在根基与偏好语义下计算得出。该流程还整合了KAOS目标建模、多层验证及符合标准的产物生成。在涵盖安全关键、金融与信息系统领域的五个案例研究中,ArgRE提供了现有框架缺失的论据级可追溯性。独立评估者对其决策理由评分显著高于启发式合成方法(4.32对比3.07,p < 0.001),表明审计能力提升;语义意图保留率仍保持在94.9%(BERTScore F1),合规覆盖率达84.7%,优于基线的47.6%–47.8%。结构分析表明,默认成对协议生成无环图,根基与偏好语义一致;跨对仲裁引入可控环路,导致两种语义产生可预测差异。
原文摘要 · Abstract (English)
As software systems grow in complexity, they must satisfy an increasing number of competing quality attributes, making it essential to balance them in a principled manner -- for example, a safety requirement for sensor-fusion verification may conflict with a tight planning-cycle budget. Multi-agent large language model frameworks support this balancing process by assigning specialized agents to different objectives. However, their conflict resolution is typically heuristic. Requirements are aggregated implicitly without explicit acceptance or rejection, limiting auditability in regulated domains. We present ArgRE, a multi-agent requirements negotiation system that embeds Dung-style abstract argumentation into the negotiation stage. Each proposal, critique, and refinement is modeled as an argument, conflicts are represented as directed attack relations, and the accepted set of arguments is computed under grounded and preferred semantics. The pipeline further integrates KAOS goal modeling, multi-layer verification, and standards-oriented artifact generation. Evaluation across five case studies spanning safety-critical, financial, and information-system domains shows that ArgRE provides argument-level traceability absent from existing frameworks. Independent evaluators rated its decision justifications significantly higher than those of heuristic synthesis (4.32 vs. 3.07, p < 0.001), indicating improved auditability, while semantic intent preservation remains comparable (94.9% BERTScore F1) and compliance coverage reaches 84.7% versus 47.6%--47.8% for baselines. Structural analysis further confirms that the default pairwise protocol yields acyclic graphs in which grounded and preferred semantics coincide, whereas cross-pair arbitration introduces controlled cyclicity, leading to predictable divergence between the two semantics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。