arXiv:2604.01529cs.AIcs.MA2026-04

用角色分工提升大模型提取健康食品政策信息的准确性

A Role-Based LLM Framework for Structured Information Extraction from Healthy Food Policies

  • 给大模型分配分析师、法律专家等角色,分步处理政策内容
  • 在608份政策上测试,显著减少幻觉和漏提关键信息
  • 适合政策研究者与公共健康数据自动化处理场景

当前大语言模型在健康食品政策信息抽取中常因文档结构多样性和不一致,出现误判、遗漏和幻觉等问题。本文提出一种基于角色的LLM框架,通过分配三种专业化角色:政策分析师负责元数据与机制分类,法律策略专家识别复杂法律手段,食品系统专家对食品系统阶段进行归类。该框架将领域知识(如法律机制定义与分类标准)嵌入角色提示,模拟真实专家分析流程。在来自健康食品政策项目(HFPP)数据库的608份政策上,使用Llama-3.3-70B模型评估,相比零样本、少样本及思维链(CoT)基线,本框架在复杂推理任务中表现更优,具备更高可靠性与可解释性,为健康政策信息自动化抽取提供有效方法。

原文摘要 · Abstract (English)

Current Large Language Model (LLM) approaches for information extraction (IE) in the healthy food policy domain are often hindered by various factors, including misinformation, specifically hallucinations, misclassifications, and omissions that result from the structural diversity and inconsistency of policy documents. To address these limitations, this study proposes a role-based LLM framework that automates the IE from unstructured policy data by assigning specialized roles: an LLM policy analyst for metadata and mechanism classification, an LLM legal strategy specialist for identifying complex legal approaches, and an LLM food system expert for categorizing food system stages. This framework mimics expert analysis workflows by incorporating structured domain knowledge, including explicit definitions of legal mechanisms and classification criteria, into role-specific prompts. We evaluate the framework using 608 healthy food policies from the Healthy Food Policy Project (HFPP) database, comparing its performance against zero-shot, few-shot, and chain-of-thought (CoT) baselines using Llama-3.3-70B. Our proposed framework demonstrates superior performance in complex reasoning tasks, offering a reliable and transparent methodology for automating IE from health policies.

信息抽取大模型应用政策分析角色分工

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。