arXiv:2508.06803cs.CLcs.MA2025-08被引 2

用多智能体协作分析讽刺,避免幻觉,准确率提升6.75%。

SEVADE: Self-Evolving Multi-Agent Analysis with Decoupled Evaluation for Hallucination-Resistant Irony Detection

  • 多智能体分工解析文本,生成结构化推理链。
  • 分离评判模块,使准确率提升6.75%,宏平均F1提高6.29%。
  • 适合需要高可靠性的讽刺检测场景,如舆情分析。

讽刺检测是自然语言处理中关键但具挑战性的任务。现有大模型方法常受限于单一视角分析、静态推理路径,以及在处理复杂讽刺修辞时易产生幻觉,影响准确性和可靠性。为此,我们提出SEVADE框架——一种基于解耦评估的自进化多智能体分析系统,用于抗幻觉讽刺检测。其核心为动态智能体推理引擎(DARE),通过基于语言学理论的多类专业智能体对文本进行多维度拆解,生成结构化推理链;随后,独立的轻量级理由裁判器(RA)仅依据该推理链完成最终分类。这种解耦架构有效降低幻觉风险。在四个基准数据集上的大量实验表明,该框架达到当前最优性能,平均准确率提升6.75%,宏平均F1提升6.29%。

原文摘要 · Abstract (English)

Sarcasm detection is a crucial yet challenging Natural Language Processing task. Existing Large Language Model methods are often limited by single-perspective analysis, static reasoning pathways, and a susceptibility to hallucination when processing complex ironic rhetoric, which impacts their accuracy and reliability. To address these challenges, we propose **SEVADE**, a novel **S**elf-**Ev**olving multi-agent **A**nalysis framework with **D**ecoupled **E**valuation for hallucination-resistant sarcasm detection. The core of our framework is a Dynamic Agentive Reasoning Engine (DARE), which utilizes a team of specialized agents grounded in linguistic theory to perform a multifaceted deconstruction of the text and generate a structured reasoning chain. Subsequently, a separate lightweight rationale adjudicator (RA) performs the final classification based solely on this reasoning chain. This decoupled architecture is designed to mitigate the risk of hallucination by separating complex reasoning from the final judgment. Extensive experiments on four benchmark datasets demonstrate that our framework achieves state-of-the-art performance, with average improvements of **6.75%** in Accuracy and **6.29%** in Macro-F1 score.

讽刺检测多智能体抗幻觉

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。