arXiv:2510.05121cs.CLcs.CE2025-10

用大模型从贸易协定中自动提取结构化知识三元组。

Towards Structured Knowledge: Advancing Triple Extraction from Regional Trade Agreements using Large Language Models

  • 用提示工程在零样本、少样本下提取法律文本中的三元组。
  • Llama 3.1 在少样本下准确率超70%,验证了方法有效性。
  • 适合法律与经济数据结构化研究者使用。

本研究探讨大型语言模型(LLMs)在经济领域提取结构化知识三元组(主语-谓语-宾语)的有效性。以区域贸易协定文本为场景,应用零样本、单样本及少样本提示技术,结合正负例输入,评估其定量与定性表现。实验采用 Llama 3.1 模型处理非结构化贸易协定文本,成功提取出贸易相关三元组信息。研究揭示了提示策略对性能的关键影响,指出当前挑战并展望未来方向,强调语言模型在经济知识图谱构建中的潜力。

原文摘要 · Abstract (English)

This study investigates the effectiveness of Large Language Models (LLMs) for the extraction of structured knowledge in the form of Subject-Predicate-Object triples. We apply the setup for the domain of Economics application. The findings can be applied to a wide range of scenarios, including the creation of economic trade knowledge graphs from natural language legal trade agreement texts. As a use case, we apply the model to regional trade agreement texts to extract trade-related information triples. In particular, we explore the zero-shot, one-shot and few-shot prompting techniques, incorporating positive and negative examples, and evaluate their performance based on quantitative and qualitative metrics. Specifically, we used Llama 3.1 model to process the unstructured regional trade agreement texts and extract triples. We discuss key insights, challenges, and potential future directions, emphasizing the significance of language models in economic applications.

知识提取大模型贸易协定三元组

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。