测试大模型在薪酬系统中的语义与语法理解能力,验证其能否准确计算并可审计。
Evaluating Semantic and Syntactic Understanding in Large Language Models for Payroll Systems
- 用分层数据集和多类提示策略评估模型对薪酬规则的理解
- 发现部分场景仅靠提示即可准确计算,复杂场景需显式计算机制
- 为高要求场景提供可复现的部署框架与实用建议
大型语言模型已广泛应用于写作、搜索与分析,其自然语言理解能力持续提升。然而,它们在精确数值计算及生成可审计输出方面仍不可靠。本文以合成薪酬系统为例,研究模型是否能理解薪酬模式、正确应用规则顺序,并实现分币级精度结果。实验涵盖从基础到复杂的分层数据集,采用从最小基线到基于模式引导与推理增强的多种提示方式,覆盖 GPT、Claude、Perplexity、Grok 与 Gemini 等多个模型系列。结果显示,在某些场景下精心设计的提示已足够,而在其他场景中则必须引入显式计算机制。该研究提出一个紧凑且可复现的评估框架,为需要高精度与可审计性的场景提供实用部署指导。
原文摘要 · Abstract (English)
Large language models are now used daily for writing, search, and analysis, and their natural language understanding continues to improve. However, they remain unreliable on exact numerical calculation and on producing outputs that are straightforward to audit. We study synthetic payroll system as a focused, high-stakes example and evaluate whether models can understand a payroll schema, apply rules in the right order, and deliver cent-accurate results. Our experiments span a tiered dataset from basic to complex cases, a spectrum of prompts from minimal baselines to schema-guided and reasoning variants, and multiple model families including GPT, Claude, Perplexity, Grok and Gemini. Results indicate clear regimes where careful prompting is sufficient and regimes where explicit computation is required. The work offers a compact, reproducible framework and practical guidance for deploying LLMs in settings that demand both accuracy and assurance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。