arXiv:2505.00455cs.HCcs.AI2025-05

用大模型帮领域专家把数据隐性知识挖出来,让可视化更懂业务。

Data Therapist: Eliciting Domain Knowledge from Subject Matter Experts Using Large Language Models

  • 通过问答+交互标注,引导专家外化数据背后的隐性知识。
  • 在会计、政治学、网络安全领域验证,发现专家思考有共性模式。
  • 适合需要结合领域经验做数据可视化的研究人员和工程师。

有效的数据可视化不仅需要技术能力,还需要对数据所处领域背景的深入理解。这种背景常包含关于数据来源、质量与使用目的的隐性知识,但通常不会明确体现在数据集中。为应对日益增长的挖掘隐性知识的需求,我们提出了Data Therapist——一个基于大语言模型的网页系统,通过混合主动性流程(迭代问答+交互式标注),帮助领域专家外化这些隐性知识。系统能自动分析用户提供的数据集,生成针对性问题,并支持多粒度标注。由此构建的结构化知识库可指导人工与自动化可视化设计。对会计、政治学和计算机安全领域的专家组合进行的定性研究表明,专家在思考数据时存在重复出现的模式,并揭示了人工智能在提升可视化设计方面的潜力。

原文摘要 · Abstract (English)

Effective data visualization requires not only technical proficiency but also a deep understanding of the domain-specific context in which data exists. This context often includes tacit knowledge about data provenance, quality, and intended use, which is rarely explicit in the dataset itself. Motivated by growing demands to surface tacit knowledge, we present the Data Therapist, a web-based system that helps domain experts externalize such implicit knowledge through a mixed-initiative process combining iterative Q&A with interactive annotation. Powered by a large language model, the system automatically analyzes user-supplied datasets, prompts users with targeted questions, and supports annotation at varying levels of granularity. The resulting structured knowledge base can inform both human and automated visualization design. A qualitative study with expert pairs from Accounting, Political Science, and Computer Security revealed recurring patterns in how expert reason about their data and highlighted opportunities for AI support to enhance visualization design.

知识挖掘大模型可视化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。