arXiv:2504.06138cs.MMcs.AI2025-04被引 3

融合多媒体与可视化分析,构建人机协同的智能洞察系统。

Multimedia and Visual Analytics in the Agentic Era

  • 提出人机协同的多媒体分析框架,整合视觉与语义理解能力。
  • 强调系统级设计而非单一算法优化,支持专业用户复杂任务。
  • 适合从事智能媒体分析、交互式数据探索的研究者与开发者。

专业用户需要工具从大规模多媒体数据中获取可操作的洞察。基础模型与AI代理的快速发展改变了这一领域,提升其准确性、可信度和推理能力是计算机视觉、机器学习与多媒体领域的热点。当前多数研究聚焦于以基准为导向的算法改进。多媒体社区应超越算法,关注完整的多媒体分析系统,支持专业用户完成复杂任务,实现真正的人机协同。利用机器学习与可视化支持用户已有数十年研究基础。本文提出一个将多媒体与可视化分析融合的框架,阐明其对现有及新型多媒体分析解决方案的影响。更多信息见 https://staff.fnwi.uva.nl/m.worring/analytics-model.html。

原文摘要 · Abstract (English)

Professional users need tools to help them gain actionable insights from large multimedia collections. Foundation models and AI agents have rapidly changed the playing field, and improving their accuracy, trustworthiness, and reasoning capabilities are active topics in the computer vision, machine learning, and multimedia communities. Most current research focuses on benchmark driven algorithmic improvements. The multimedia community is the place to go beyond algorithms and consider complete multimedia analytics systems that support professional users in their complex tasks and achieve a true teaming of humans and AI. Supporting users with machine learning and visualizations has been studied for decades in the visual analytics field. In this paper, we propose a framework to bring multimedia and visual analytics together and indicate how it could impact current and new multimedia analytics solutions. Additional information can be found at https://staff.fnwi.uva.nl/m.worring/analytics-model.html

多媒体分析人机协同可视化AI代理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。