arXiv:2508.11120cs.CL2025-08被引 1

用反思、记忆和规划提升营销多智能体系统的可靠性

Towards Reliable Multi-Agent Systems for Marketing Applications via Reflection, Memory, and Planning

  • 引入RAMP框架,通过迭代计划、工具调用与验证生成高质量受众
  • 在88个查询上准确率提升28个百分点,模糊任务召回率提高约20%
  • 适合需要高可靠性的企业级AI应用开发者参考

大语言模型(LLM)的发展使得能够规划并调用工具完成复杂任务的AI智能体成为可能,但其在真实场景中的可靠性研究仍有限。本文针对营销任务中的受众筛选问题,提出RAMP框架:通过迭代式规划、工具调用、输出验证及改进建议,持续优化结果。同时引入长期记忆存储,保留客户特定事实与历史查询。实验表明,在88个评估查询中,该方法使准确率提升28个百分点;在较小挑战集上,每增加一次验证与反思迭代,召回率平均提升约20个百分点,并显著提高用户满意度。结果为面向动态产业环境的可靠LLM系统部署提供了实践指导。

原文摘要 · Abstract (English)

Recent advances in large language models (LLMs) enabled the development of AI agents that can plan and interact with tools to complete complex tasks. However, literature on their reliability in real-world applications remains limited. In this paper, we introduce a multi-agent framework for a marketing task: audience curation. To solve this, we introduce a framework called RAMP that iteratively plans, calls tools, verifies the output, and generates suggestions to improve the quality of the audience generated. Additionally, we equip the model with a long-term memory store, which is a knowledge base of client-specific facts and past queries. Overall, we demonstrate the use of LLM planning and memory, which increases accuracy by 28 percentage points on a set of 88 evaluation queries. Moreover, we show the impact of iterative verification and reflection on more ambiguous queries, showing progressively better recall (roughly +20 percentage points) with more verify/reflect iterations on a smaller challenge set, and higher user satisfaction. Our results provide practical insights for deploying reliable LLM-based systems in dynamic, industry-facing environments.

多智能体营销应用LLM可靠性记忆机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。