arXiv:2511.11821cs.CLcs.AI2025-11

70B以下大模型在水电监管信息提取中存在14B性能拐点,小模型易幻觉。

Scaling Open-Weight Large Language Models for Hydropower Regulatory Information Extraction: A Systematic Analysis

  • 测试0.6B-70B参数的7个开源模型,发现14B是性能分水岭。
  • 14B以上模型F1达64%,超70B模型逼近77%,小模型仅51%。
  • 揭示小模型高召回即失败的系统性幻觉问题,适合合规部署参考。

利用大语言模型从监管文件中提取信息,在性能与计算资源间存在关键权衡。我们评估了7个开源模型(0.6B-70B参数)在水电许可文档上的表现,提供实证部署指导。分析发现,14B参数阈值处验证方法发生显著转变:低于此阈值时无效(F1 < 0.15),高于则可行(F1 = 0.64)。消费级可部署模型通过适当验证达到64% F1,小模型性能停滞于51%;大规模模型接近77% F1,但需企业级基础设施。识别出系统性幻觉模式:小模型中完美召回实为提取失败而非成功。研究建立首个面向开放权重模型在监管场景下的资源-性能映射,支持基于证据的模型选择,对水电合规具有即时价值,并为信息抽取任务中的参数缩放效应提供普适性洞见。

原文摘要 · Abstract (English)

Information extraction from regulatory documents using large language models presents critical trade-offs between performance and computational resources. We evaluated seven open-weight models (0.6B-70B parameters) on hydropower licensing documentation to provide empirical deployment guidance. Our analysis identified a pronounced 14B parameter threshold where validation methods transition from ineffective (F1 $<$ 0.15) to viable (F1 = 0.64). Consumer-deployable models achieve 64\% F1 through appropriate validation, while smaller models plateau at 51\%. Large-scale models approach 77\% F1 but require enterprise infrastructure. We identified systematic hallucination patterns where perfect recall indicates extraction failure rather than success in smaller models. Our findings establish the first comprehensive resource-performance mapping for open-weight information extraction in regulatory contexts, enabling evidence-based model selection. These results provide immediate value for hydropower compliance while contributing insights into parameter scaling effects that generalize across information extraction tasks.

信息提取大模型水电合规参数规模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。