arXiv:2510.14184cs.LGcs.AI2025-10AAAI

MAFA用可配置多智能体系统解决企业级标注难题,大幅减少人工工作量。

MAFA: A Multi-Agent Framework for Enterprise-Scale Annotation with Configurable Task Adaptation

  • 通过可配置的多智能体协作与裁判共识机制实现动态任务适配
  • 消除百万级语音标注积压,人工标注节省超5000小时/年
  • 适合金融等行业需高精度、可扩展标注的场景

我们提出MAFA(多智能体标注框架),一个已投入生产的系统,通过可配置的多智能体协作重构企业级标注流程。针对金融服务业中数百万客户语句需精准分类的痛点,MAFA结合专用智能体、结构化推理与基于裁判的共识机制。其独特支持动态任务适配,可通过配置而非代码修改定义自定义标注类型(如FAQ、意图、实体或领域特定类别)。在摩根大通部署后,该系统成功消除100万条语句的标注积压,平均与人工标注者达成86%的一致性,每年节省超5000小时手动标注时间。系统对语句进行置信度分类,各类别占比分别为:高置信度85%、中等10%、低5%,使人工仅需处理模糊及低覆盖情况。我们在多个数据集和语言上验证了其有效性,相比传统与单智能体基线,内部意图分类数据集上提升13.8%的Top-1准确率、15.1%的Top-5准确率、16.9%的F1值,公共基准测试亦有类似提升。本工作弥合了理论多智能体系统与实际企业部署间的差距,为面临相似标注挑战的组织提供可复制蓝图。

原文摘要 · Abstract (English)

We present MAFA (Multi-Agent Framework for Annotation), a production-deployed system that transforms enterprise-scale annotation workflows through configurable multi-agent collaboration. Addressing the critical challenge of annotation backlogs in financial services, where millions of customer utterances require accurate categorization, MAFA combines specialized agents with structured reasoning and a judge-based consensus mechanism. Our framework uniquely supports dynamic task adaptation, allowing organizations to define custom annotation types (FAQs, intents, entities, or domain-specific categories) through configuration rather than code changes. Deployed at JP Morgan Chase, MAFA has eliminated a 1 million utterance backlog while achieving, on average, 86% agreement with human annotators, annually saving over 5,000 hours of manual annotation work. The system processes utterances with annotation confidence classifications, which are typically 85% high, 10% medium, and 5% low across all datasets we tested. This enables human annotators to focus exclusively on ambiguous and low-coverage cases. We demonstrate MAFA's effectiveness across multiple datasets and languages, showing consistent improvements over traditional and single-agent annotation baselines: 13.8% higher Top-1 accuracy, 15.1% improvement in Top-5 accuracy, and 16.9% better F1 in our internal intent classification dataset and similar gains on public benchmarks. This work bridges the gap between theoretical multi-agent systems and practical enterprise deployment, providing a blueprint for organizations facing similar annotation challenges.

多智能体标注系统金融AI自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。