arXiv:2603.05832cs.HCcs.AI2026-03被引 2

为对话式可视化分析设计了无需编程的LLM评估工具。

Lexara: A User-Centered Toolkit for Evaluating Large Language Models for Conversational Visual Analytics

  • 基于22名开发者和16名用户调研,构建真实场景测试用例。
  • 融合规则与LLM评判,量化评估图表与文本输出质量。
  • 交互式界面支持无代码实验,适合研究者与工程师使用。

大型语言模型(LLMs)正推动对话式可视化分析(CVA)发展,实现通过自然语言进行数据分析。然而,评估LLM在CVA中的表现仍面临挑战:需编程技能、忽略真实场景复杂性、缺乏对多格式输出(图表与文本)的可解释评估指标。通过对22名CVA开发者和16名终端用户访谈,我们识别出典型应用场景、评估标准与工作流程。提出Lexara——一个以用户为中心的CVA评估工具包,包含:(i) 覆盖真实场景的测试用例;(ii) 可解释的评估指标,涵盖可视化质量(数据保真度、语义对齐、功能正确性、设计清晰度)与语言质量(事实准确性、分析推理、对话连贯性),采用规则方法与LLM-as-a-Judge相结合;(iii) 交互式工具,支持无编程的实验配置与多格式、多层次结果探索。通过为期两周的日记研究,六名来自初始22人的开发者反馈表明,Lexara有效辅助模型与提示选择。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are transforming Conversational Visual Analytics (CVA) by enabling data analysis through natural language. However, evaluating LLMs for CVA remains a challenge: requiring programming expertise, overlooking real-world complexity, and lacking interpretable metrics for multi-format (visualizations and text) outputs. Through interviews with 22 CVA developers and 16 end-users, we identified use cases, evaluation criteria and workflows. We present Lexara, a user-centered evaluation toolkit for CVA that operationalizes these insights into: (i) test cases spanning real-world scenarios; (ii) interpretable metrics covering visualization quality (data fidelity, semantic alignment, functional correctness, design clarity) and language quality (factual grounding, analytical reasoning, conversational coherence) using rule-based and LLM-as-a-Judge methods; and (iii) an interactive toolkit enabling experimental setup and multi-format and multi-level exploration of results without programming expertise. We conducted a two-week diary study with six CVA developers, drawn from our initial cohort of 22. Their feedback demonstrated Lexara's effectiveness for guiding appropriate model and prompt selection.

对话式分析LLM评估可视化工具包

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。