用大模型自动构建带时间戳的细粒度观点知识库
Towards Temporal Knowledge-Base Creation for Fine-Grained Opinion Analysis with Language Models
- 将主流观点挖掘方法融入声明式大模型标注流程
- 在测试集上实现与人工标注相当的标签一致性
- 适合需要时间敏感观点分析的研究者使用
我们提出一种可扩展的方法,利用大语言模型(LLMs)作为自动化标注器,构建时序观点知识库。尽管时序文本观点分析在预测和趋势分析等下游任务中具有潜力,但现有方法因缺乏时间锚定的细粒度标注而未充分挖掘该潜力。本方法通过将已确立的观点挖掘范式整合进声明式大模型标注流水线,实现无需手动提示工程的结构化观点抽取。定义了三个基于情感与观点挖掘文献的数据模型,作为结构化表示的模式。使用人工标注样本进行严格定量评估,并采用两个独立的LLM完成最终标注,按细粒度观点维度计算标注者间一致性,类似人类标注协议。生成的知识库包含时间对齐的结构化观点,兼容检索增强生成(RAG)、时序问答和时间线摘要等应用。
原文摘要 · Abstract (English)
We propose a scalable method for constructing a temporal opinion knowledge base with large language models (LLMs) as automated annotators. Despite the demonstrated utility of time-series opinion analysis of text for downstream applications such as forecasting and trend analysis, existing methodologies underexploit this potential due to the absence of temporally grounded fine-grained annotations. Our approach addresses this gap by integrating well-established opinion mining formulations into a declarative LLM annotation pipeline, enabling structured opinion extraction without manual prompt engineering. We define three data models grounded in sentiment and opinion mining literature, serving as schemas for structured representation. We perform rigorous quantitative evaluation of our pipeline using human-annotated test samples. We carry out the final annotations using two separate LLMs, and inter-annotator agreement is computed label-wise across the fine-grained opinion dimensions, analogous to human annotation protocols. The resulting knowledge base encapsulates time-aligned, structured opinions and is compatible with applications in Retrieval-Augmented Generation (RAG), temporal question answering, and timeline summarisation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。