arXiv:2410.00727cs.LGcs.CL2024-10

用图文结合工具帮分析师快速发现多维数据中的异常

"Show Me What's Wrong!": Combining Charts and Text to Guide Data Analysis

  • 按分析区域分割数据,自动标记需关注部分
  • 选中区域后生成文本摘要和图形展示
  • 适合金融反欺诈等需要快速定位异常的场景

在多维数据中分析并发现异常是一项繁琐但至关重要的任务,尤其在金融欺诈检测中,分析师需从交易数据中快速识别可疑行为。这一过程涉及模式识别、分组与比较等复杂探索性操作。为缓解信息过载问题,我们提出一种结合自动化信息高亮、大语言模型生成文本洞察与可视化分析的工具,支持不同粒度的探索。系统对数据按分析区域进行划分,并以可视化方式呈现,利用自动视觉提示标识需重点关注的部分。用户选定某区域后,系统提供文本与图形双重摘要,文本作为高层与细节视图间的桥梁,帮助快速理解关键信息;进一步的详细探索可通过图形化表示完成。七位领域专家参与的反馈研究显示,该工具能有效支持并引导探索性分析,显著降低可疑信息的识别难度。

原文摘要 · Abstract (English)

Analyzing and finding anomalies in multi-dimensional datasets is a cumbersome but vital task across different domains. In the context of financial fraud detection, analysts must quickly identify suspicious activity among transactional data. This is an iterative process made of complex exploratory tasks such as recognizing patterns, grouping, and comparing. To mitigate the information overload inherent to these steps, we present a tool combining automated information highlights, Large Language Model generated textual insights, and visual analytics, facilitating exploration at different levels of detail. We perform a segmentation of the data per analysis area and visually represent each one, making use of automated visual cues to signal which require more attention. Upon user selection of an area, our system provides textual and graphical summaries. The text, acting as a link between the high-level and detailed views of the chosen segment, allows for a quick understanding of relevant details. A thorough exploration of the data comprising the selection can be done through graphical representations. The feedback gathered in a study performed with seven domain experts suggests our tool effectively supports and guides exploratory analysis, easing the identification of suspicious information.

数据可视化异常检测LLM应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。