arXiv:2509.14943cs.CL2025-09

用隐含文本测试大模型信息抽取能力,发现微调可提升隐性推理表现。

Explicit vs. Implicit Biographies: Evaluating and Adapting LLM Information Extraction on Wikidata-Derived Texts

  • 构建10k条隐含与显式传记文本数据集,对比大模型表现
  • 微调后模型在隐含文本上准确率显著提升,尤其在LoRA适配下
  • 适合关注大模型推理能力与可解释性的研究人员

文本隐含性一直是自然语言处理中的难点,传统方法依赖显式陈述识别实体及其关系。例如‘Zuhdi每周日参加教堂活动’对人类而言可推断其与基督教的关系,但自动推理仍具挑战。大语言模型(LLMs)在文本理解与信息抽取(IE)等下游任务中表现优异。本研究考察文本隐含性对预训练LLM(LLaMA 2.3、DeepSeekV1、Phi1.5)信息抽取的影响,生成两组各10,000条合成数据,分别对应传记信息的隐含与显式表述。通过实验评估模型性能,并分析在隐含数据上微调是否能提升其泛化能力。结果表明,使用LoRA(低秩适配)进行微调可显著增强模型从隐含文本中提取信息的能力,提升了模型的可解释性与可靠性。

原文摘要 · Abstract (English)

Text Implicitness has always been challenging in Natural Language Processing (NLP), with traditional methods relying on explicit statements to identify entities and their relationships. From the sentence "Zuhdi attends church every Sunday", the relationship between Zuhdi and Christianity is evident for a human reader, but it presents a challenge when it must be inferred automatically. Large language models (LLMs) have proven effective in NLP downstream tasks such as text comprehension and information extraction (IE). This study examines how textual implicitness affects IE tasks in pre-trained LLMs: LLaMA 2.3, DeepSeekV1, and Phi1.5. We generate two synthetic datasets of 10k implicit and explicit verbalization of biographic information to measure the impact on LLM performance and analyze whether fine-tuning implicit data improves their ability to generalize in implicit reasoning tasks. This research presents an experiment on the internal reasoning processes of LLMs in IE, particularly in dealing with implicit and explicit contexts. The results demonstrate that fine-tuning LLM models with LoRA (low-rank adaptation) improves their performance in extracting information from implicit texts, contributing to better model interpretability and reliability.

信息抽取大模型隐含推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。