arXiv:2603.18010cs.CLcs.AI2026-03

用AI自动提取政治人物多维履历,突破人工耗时瓶颈。

Agentic Framework for Political Biography Extraction

  • 分两阶段:先用智能体聚合网络信息,再结构化编码
  • 模型准确率媲美甚至超过人类专家,且挖掘信息量超维基百科
  • 能缓解长文本和多语言导致的偏差,适合政治科学数据构建

大规模政治数据集的构建通常需从海量非结构化文档或网络来源中提取结构化事实,传统方法依赖昂贵的人力专家,难以规模化自动化。本文利用大语言模型(LLMs)实现多维度精英人物履历的自动化提取,解决政治科学中的长期瓶颈。提出“合成-编码”两阶段框架:上游合成阶段采用递归智能体LLM从异构网络源中搜索、筛选并整理传记内容;下游编码阶段将整理后的内容映射为结构化数据框。通过三项主要验证:其一,在提供整理上下文的前提下,LLM编码器在提取准确率上达到或超过人类专家水平;其二,在网络环境中,该智能体系统从资源中获取的信息量超过集体智慧(维基百科);其三,直接对长篇多语言语料进行编码会引入偏差,而合成阶段通过提炼为高信噪比表示可有效缓解该问题。经全面评估,本研究提供了一种可泛化、可扩展的透明化大规模政治数据库构建框架。

原文摘要 · Abstract (English)

The production of large-scale political datasets typically demands extracting structured facts from vast piles of unstructured documents or web sources, a task that traditionally relies on expensive human experts and remains prohibitively difficult to automate at scale. In this paper, we leverage Large Language Models (LLMs) to automate the extraction of multi-dimensional elite biographies, addressing a long-standing bottleneck in political science research. We propose a two-stage ``Synthesis-Coding'' framework for complex extraction task: an upstream synthesis stage that uses recursive agentic LLMs to search, filter, and curate biography from heterogeneous web sources, followed by a downstream coding stage that maps curated biography into structured dataframes. We validate this framework through three primary results. First, we demonstrate that, when given curated contexts, LLM coders match or outperform human experts in extraction accuracy. Second, we show that in web environments, the agentic system synthesizes more information from web resources than human collective intelligence (Wikipedia). Finally, we diagnosed that directly coding from long and multi-language corpora introduces bias that the synthesis stage can alleviate by curating evidence into signal-dense representations. By comprehensive evaluation, We provide a generalizable, scalable framework for building transparent and expansible large scale database in political science.

知识抽取智能体政治数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。