arXiv:2510.00245cs.HCcs.AI2025-10被引 2

测试AI理解会议中关于图表对话的能力,发现纯文本效果最佳。

Can AI agents understand spoken conversations about data visualizations in online meetings?

  • 用双轴框架评估AI对图表对话的理解能力
  • 纯文本输入在72段对话中达96%准确率
  • 适合研究会议AI助手或多模态理解的学者

本文评估AI代理在在线会议场景中理解关于数据可视化口头对话的能力。随着会议辅助AI的发展,其服务质量取决于模型对对话的理解程度。为此,我们提出一种双轴测试框架,用于诊断AI对数据对话的领会情况。基于该框架,设计了一系列测试,评估一个包含72段关于数据可视化的口语对话的新语料库。我们考察了多种管道和模型架构(LLM与VLM),以及不同可视化输入形式(图表图像、源代码或两者结合)对模型表现的影响。结果显示,在我们的测试中,仅使用文本输入的模态取得了最佳性能(96%)。

原文摘要 · Abstract (English)

In this short paper, we present work evaluating an AI agent's understanding of spoken conversations about data visualizations in an online meeting scenario. There is growing interest in the development of AI-assistants that support meetings, such as by providing assistance with tasks or summarizing a discussion. The quality of this support depends on a model that understands the conversational dialogue. To evaluate this understanding, we introduce a dual-axis testing framework for diagnosing the AI agent's comprehension of spoken conversations about data. Using this framework, we designed a series of tests to evaluate understanding of a novel corpus of 72 spoken conversational dialogues about data visualizations. We examine diverse pipelines and model architectures, LLM vs VLM, and diverse input formats for visualizations (the chart image, its underlying source code, or a hybrid of both) to see how this affects model performance on our tests. Using our evaluation methods, we found that text-only input modalities achieved the best performance (96%) in understanding discussions of visualizations in online meetings.

AI助手多模态理解会议分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。