arXiv:2504.00343cs.CL2025-04被引 3

用大模型自动提取学术文献中的概念定义,聚焦媒体偏见领域。

Leveraging Large Language Models for Automated Definition Extraction with TaxoMatic A Case Study on Media Bias

  • 基于大模型构建自动定义抽取框架,整合文献筛选与定义提取。
  • 在2398篇人工标注文章上验证,Claude-3-sonnet表现最佳。
  • 适合需要快速构建领域知识体系的研究者使用。

本文提出TaxoMatic框架,利用大语言模型自动化从学术文献中提取概念定义。聚焦媒体偏见领域,该框架包含数据收集、基于LLM的相关性分类及概念定义提取。在包含2,398篇人工标注文章的数据集上进行评估,结果显示Claude-3-sonnet在相关性分类和定义提取任务中均表现最优。未来工作包括扩展数据集并将其应用于其他领域。

原文摘要 · Abstract (English)

This paper introduces TaxoMatic, a framework that leverages large language models to automate definition extraction from academic literature. Focusing on the media bias domain, the framework encompasses data collection, LLM-based relevance classification, and extraction of conceptual definitions. Evaluated on a dataset of 2,398 manually rated articles, the study demonstrates the frameworks effectiveness, with Claude-3-sonnet achieving the best results in both relevance classification and definition extraction. Future directions include expanding datasets and applying TaxoMatic to additional domains.

大模型定义提取媒体偏见

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。