让数字人文学者交互式聚类文档,快速发现主题与情感。
Perspectives - Interactive Document Clustering in the Discourse Analysis Tool Suite
- 通过用户提示重写文档并生成嵌入向量,聚焦特定分析视角。
- 支持人工干预调整聚类结果,可微调嵌入模型以匹配研究目标。
- 适合需深入文本分析的数字人文研究者,提升数据探索效率。
本文介绍Perspectives,作为话语分析工具套件的交互式扩展,旨在赋能数字人文(DH)学者探索和组织大规模非结构化文档集合。Perspectives实现了一种灵活、面向特定方面的文档聚类流程,并具备人机协同优化能力。通过文档重写提示和基于指令的嵌入向量,研究者可初步设定分析视角;随后借助聚类优化工具与嵌入模型微调机制,进一步对齐用户意图。演示展示了典型工作流:研究者利用Perspectives的交互式文档地图,识别主题、情感或其他相关类别,从而获得洞察并为后续深度分析准备数据。
原文摘要 · Abstract (English)
This paper introduces Perspectives, an interactive extension of the Discourse Analysis Tool Suite designed to empower Digital Humanities (DH) scholars to explore and organize large, unstructured document collections. Perspectives implements a flexible, aspect-focused document clustering pipeline with human-in-the-loop refinement capabilities. We showcase how this process can be initially steered by defining analytical lenses through document rewriting prompts and instruction-based embeddings, and further aligned with user intent through tools for refining clusters and mechanisms for fine-tuning the embedding model. The demonstration highlights a typical workflow, illustrating how DH researchers can leverage Perspectives's interactive document map to uncover topics, sentiments, or other relevant categories, thereby gaining insights and preparing their data for subsequent in-depth analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。