测试大模型在金融合规中的规则执行能力,发现仅靠规则引用仍会违规。
ReguSim: Evaluating LLM Agent Rule Grounding in Financial Compliance

- 构建金融合规环境与基准测试,分离推理、行为、执行和监控四类证据。
- 可见规则可减少违规,但无法杜绝;激励或角色设定会改变行为模式。
- 模型推理常误导监控者,需展示执行证据才能准确判断合规性。
金融领域的大型语言模型代理可能引用规则却仍提交违反可执行约束的订单,或误读监管监控证据。本文提出 ReguSim 控制型金融合规环境与 ReguBench 监控基准,将四个关键环节分离:声明推理、尝试动作、执行约束、监控证据。在 DeepSeek V4 Pro 和 Gemini 3.5 Flash 的交易员实验中,可见规则虽能降低但无法消除被拒操作;激励或角色设定显著影响行为。桥接研究显示,若不提供执行证据,独立监控者易被模型推理误导。在监控任务中,简单结构化基线方法表现不逊于纯提示式 LLM。结果表明,金融合规评估应聚焦规则与行动、证据的对齐,而非单一合规评分。
原文摘要 · Abstract (English)
LLM agents in financial markets may cite rules yet still submit orders that violate executable constraints or misread surveillance evidence. We introduce ReguSim, a controlled financial-compliance environment, and ReguBench, a target-marked monitoring benchmark, to separate four artifacts: stated reasoning, attempted action, execution enforcement, and monitor evidence. In trader runs with DeepSeek V4 Pro and Gemini 3.5 Flash, visible rules reduce but do not eliminate rejected actions, and incentive or persona framing shifts behavior. A bridge study shows that trader rationales can mislead an independent monitor unless enforcement evidence is shown. In monitoring, simple structured baselines either match or exceed prompt-only LLMs. The results frame financial compliance evaluation as an audit of rule-grounded actions and evidence use, rather than a single compliance score.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。