构建面向印度零售银行的可控对话系统,提升安全与准确性。
FiMI Banking: A Sovereign Model for Indian Retail Banking
- 基于权威银行文档与合成客户数据构建专用对话环境。
- 偏好优化使无关请求拒绝率从52%升至80%,强化学习提升任务准确率至0.718。
- 适合关注金融场景安全、合规及工具调用的开发者与研究者。
银行需要能够回答产品问题、协助账户操作,并在严格运营与监管约束下安全运行的对话系统。通用语言模型难以满足这些要求,尤其在需基于事实、正确使用工具或谨慎处理银行敏感情境时表现不足。本文提出FiMI Banking,一个面向印度零售银行的受控场景。该系统基于经审核的银行文档、结构化真实答案、合成客户背景及银行工具构建。我们评估两种后训练方法:用于响应行为的偏好优化,以及基于可验证奖励的强化学习,用于多轮工具使用任务。偏好优化显著提升安全性:无关请求拒绝率从52%提高至80%。强化学习将边缘情况性能从0.509提升至0.718,顺序敏感任务性能从0.590提升至0.679,同时生成令牌减少29%。结果表明,偏好优化与可验证奖励强化学习分别解决可靠银行代理的不同核心需求。
原文摘要 · Abstract (English)
Banks need conversational systems that can answer product questions, assist customers with account-related requests, and operate safely within strict operational and regulatory constraints. General-purpose language models do not reliably meet these requirements. They fall short when a task requires grounded information, correct tool use, or cautious handling of bank-specific sensitive situations. We introduce FiMI Banking, a controlled Indian retail-banking setting. We build it from vetted banking documents, structured ground truth, synthetic customer backgrounds, and banking tools. We evaluate two post-training approaches: preference optimization for response-level behavior, and reinforcement learning with verifiable rewards for multi-turn tool-use tasks. Preference optimization improves safe behavior substantially: out-of-scope refusal rises from 52% to 80%. Reinforcement learning improves edge-case performance from 0.509 to 0.718 and order-sensitive task performance from 0.590 to 0.679, while using 29% fewer generated tokens. These results show that preference optimization and verifiable-reward reinforcement learning address complementary requirements for reliable banking agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。