用自然语言对话访问人文数据库,让研究更高效。
Talking to Data: Designing Smart Assistants for Humanities Databases
- 基于大模型与RAG技术,支持自然语言查询
- 在俄罗斯20世纪日记档案上测试,响应质量高
- 适合历史人类学研究者和普通用户使用
人文研究数据库的访问常受限于传统交互方式,尤其在搜索与响应生成方面。本研究提出一种基于大语言模型的智能助手,以聊天机器人形式实现与数字人文数据的自然语言交互。该助手采用RAG框架,集成混合搜索、自动查询生成、文本转SQL过滤、语义数据库检索和超链接插入等技术。通过在包含20世纪俄语使用者日记的Prozhito数字档案上进行实验,评估了多种语言模型的响应质量。系统专为人类学与历史研究者及无技术背景的公众设计,无需预先培训即可使用。该工具使研究人员能通过自然语言查询复杂数据库,提升人文研究的可及性与效率。研究展示了大模型在改变公众与数字档案互动方式上的潜力,使交互更直观、包容。补充材料可在GitHub仓库获取:https://github.com/alekosus/talking-to-data-intersys2025。
原文摘要 · Abstract (English)
Access to humanities research databases is often hindered by the limitations of traditional interaction formats, particularly in the methods of searching and response generation. This study introduces an LLM-based smart assistant designed to facilitate natural language communication with digital humanities data. The assistant, developed in a chatbot format, leverages the RAG approach and integrates state-of-the-art technologies such as hybrid search, automatic query generation, text-to-SQL filtering, semantic database search, and hyperlink insertion. To evaluate the effectiveness of the system, experiments were conducted to assess the response quality of various language models. The testing was based on the Prozhito digital archive, which contains diary entries from predominantly Russian-speaking individuals who lived in the 20th century. The chatbot is tailored to support anthropology and history researchers, as well as non-specialist users with an interest in the field, without requiring prior technical training. By enabling researchers to query complex databases with natural language, this tool aims to enhance accessibility and efficiency in humanities research. The study highlights the potential of Large Language Models to transform the way researchers and the public interact with digital archives, making them more intuitive and inclusive. Additional materials are presented in GitHub repository: https://github.com/alekosus/talking-to-data-intersys2025.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。