用多个智能体协作分析反语,提升大模型识别准确率。
CAF-I: A Collaborative Multi-Agent Framework for Enhanced Irony Detection with Large Language Models
- 设计上下文、语义、修辞三类智能体协同分析
- 零样本测试下宏平均F1达76.31,提升4.98个百分点
- 适合需要高精度与可解释性的反语检测场景
大语言模型已成为反语检测的主流方法,但现有方法存在单一视角局限、理解不全面和缺乏可解释性等问题。本文提出协同反语检测框架CAF-I,基于大语言模型构建多智能体系统,包含上下文、语义和修辞三类专业智能体,进行多维度分析并交互优化。决策智能体整合各视角意见,精炼评估智能体提供条件反馈以持续优化。在基准数据集上的实验表明,CAF-I实现当前最优零样本性能,多数指标达到领先水平,平均宏平均F1为76.31,相比最强基线绝对提升4.98。该成果得益于对人类多视角分析的有效模拟,显著提升了检测准确率与可解释性。
原文摘要 · Abstract (English)
Large language model (LLM) have become mainstream methods in the field of sarcasm detection. However, existing LLM methods face challenges in irony detection, including: 1. single-perspective limitations, 2. insufficient comprehensive understanding, and 3. lack of interpretability. This paper introduces the Collaborative Agent Framework for Irony (CAF-I), an LLM-driven multi-agent system designed to overcome these issues. CAF-I employs specialized agents for Context, Semantics, and Rhetoric, which perform multidimensional analysis and engage in interactive collaborative optimization. A Decision Agent then consolidates these perspectives, with a Refinement Evaluator Agent providing conditional feedback for optimization. Experiments on benchmark datasets establish CAF-I's state-of-the-art zero-shot performance. Achieving SOTA on the vast majority of metrics, CAF-I reaches an average Macro-F1 of 76.31, a 4.98 absolute improvement over the strongest prior baseline. This success is attained by its effective simulation of human-like multi-perspective analysis, enhancing detection accuracy and interpretability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。