arXiv:2412.18428cs.AIcs.CL2024-12中稿 · the IJCNLP AACL 20…被引 7

用语言代理统一查询数据库与图文数据,提升多模态探索效率。

Multi-Modal Data Exploration via Language Agents

  • 通过语言代理分解自然语言问题,协调文本、图像等专家协同处理。
  • 在多模态数据集上表现优于现有系统,准确率和响应速度均提升。
  • 适合需要跨数据类型快速分析的科研与企业用户。

国际企业、组织和医院收集了大量存储在数据库、文本文档、图像和视频中的多模态数据。尽管在多模态数据探索和自然语言转数据库查询语言方面已有进展,但如何用自然语言同时查询结构化数据库和非结构化模态(如文本、图像)仍是未被充分研究的挑战。本文提出M²EX——一个基于语言代理的多模态数据探索系统。该系统受真实应用场景启发,利用基于大模型的代理框架,将自然语言问题分解为子任务(如文本转SQL、图像分析),并高效编排特定模态专家执行查询计划。在包含关系型数据、文本和图像的多模态数据集上的实验表明,该系统在准确率、查询延迟、API成本和规划效率等多项指标上均优于现有先进系统,得益于对大模型推理能力的更优利用。

原文摘要 · Abstract (English)

International enterprises, organizations, and hospitals collect large amounts of multi-modal data stored in databases, text documents, images, and videos. While there has been recent progress in the separate fields of multi-modal data exploration as well as in database systems that automatically translate natural language questions to database query languages, the research challenge of querying both structured databases and unstructured modalities (e.g., texts, images) in natural language remains largely unexplored. In this paper, we propose M$^2$EX -a system that enables multi-modal data exploration via language agents. Our approach is based on the following research contributions: (1) Our system is inspired by a real-world use case that enables users to explore multi-modal information systems. (2) M$^2$EX leverages an LLM-based agentic AI framework to decompose a natural language question into subtasks such as text-to-SQL generation and image analysis and to orchestrate modality-specific experts in an efficient query plan. (3) Experimental results on multi-modal datasets, encompassing relational data, text, and images, demonstrate that our system outperforms state-of-the-art multi-modal exploration systems, excelling in both accuracy and various performance metrics, including query latency, API costs, and planning efficiency, thanks to the more effective utilization of the reasoning capabilities of LLMs.

多模态语言代理自然语言查询AI系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。