arXiv:2501.00906cs.MAcs.AI2025-01被引 16

用大模型+多智能体实现多媒体物联网的自动事件处理。

Large Language Model Based Multi-Agent System Augmented Complex Event Processing Pipeline for Internet of Multimedia Things

  • 基于Autogen和Kafka构建多智能体协同处理流水线。
  • 高复杂度视频下延迟上升,但叙事一致性保持稳定。
  • 适合研究AI系统集成与实时多媒体分析的开发者。

本文提出一种基于大型语言模型(LLM)的多智能体系统框架,用于复杂事件处理(CEP),聚焦视频查询处理场景。目标是构建一个原型系统,将前沿的LLM编排框架与发布/订阅(pub/sub)工具结合,解决LLM与现有CEP系统集成问题。采用Autogen框架配合Kafka消息代理,实现了可自主运行的CEP流水线,能处理复杂工作流。通过大量实验评估不同配置、复杂度及视频分辨率下的性能表现,揭示了功能与延迟之间的权衡。结果显示,智能体数量增加或视频复杂度提高会带来延迟上升,但系统在叙事连贯性方面保持高度一致。该研究拓展了分布式AI系统的创新方法,为将此类系统集成到现有基础设施提供了详细见解。

原文摘要 · Abstract (English)

This paper presents the development and evaluation of a Large Language Model (LLM), also known as foundation models, based multi-agent system framework for complex event processing (CEP) with a focus on video query processing use cases. The primary goal is to create a proof-of-concept (POC) that integrates state-of-the-art LLM orchestration frameworks with publish/subscribe (pub/sub) tools to address the integration of LLMs with current CEP systems. Utilizing the Autogen framework in conjunction with Kafka message brokers, the system demonstrates an autonomous CEP pipeline capable of handling complex workflows. Extensive experiments evaluate the system's performance across varying configurations, complexities, and video resolutions, revealing the trade-offs between functionality and latency. The results show that while higher agent count and video complexities increase latency, the system maintains high consistency in narrative coherence. This research builds upon and contributes to, existing novel approaches to distributed AI systems, offering detailed insights into integrating such systems into existing infrastructures.

多智能体视频处理大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。