arXiv:2508.01612cs.LGcs.AI2025-08被引 1

引入外部代理提升模型决策能力,让人类与机器协同优化学习过程。

Augmented Reinforcement Learning Framework For Enhancing Decision-Making In Machine Learning Models Using External Agents

  • 用两个外部代理实时评估并修正模型决策路径。
  • 结合人工反馈后,模型在复杂场景下准确率显著提升。
  • 适合需要高可靠性、人机协作的金融等数据敏感领域。

本文提出一种增强型强化学习框架(ARL),用于提升机器学习模型的决策能力。通过引入外部代理作为监督者,可对模型决策进行校正。外部代理1为实时评估者,识别次优动作并形成被拒数据流;外部代理2则对反馈进行有选择性地筛选,确保业务场景下的相关性和准确性,生成可用于后续训练的核准数据集。该框架在“文档识别与信息提取”这一实际应用场景中验证,该问题主要源于银行系统,但可扩展至其他领域。实验表明,融入人工反馈后,模型在复杂或模糊环境中的鲁棒性和决策准确性明显提高。结合机器效率与人类洞察力的增强方法,实现了更高学习标准,证明了此类人机协同强化学习框架在数据驱动应用中的可扩展性与有效性。

原文摘要 · Abstract (English)

This work proposes a novel technique Augmented Reinforcement Learning framework for the improvement of decision-making capabilities of machine learning models. The introduction of agents as external overseers checks on model decisions. The external agent can be anyone, like humans or automated scripts, that helps in decision path correction. It seeks to ascertain the priority of the "Garbage-In, Garbage-Out" problem that caused poor data inputs or incorrect actions in reinforcement learning. The ARL framework incorporates two external agents that aid in course correction and the guarantee of quality data at all points of the training cycle. The External Agent 1 is a real-time evaluator, which will provide feedback light of decisions taken by the model, identify suboptimal actions forming the Rejected Data Pipeline. The External Agent 2 helps in selective curation of the provided feedback with relevance and accuracy in business scenarios creates an approved dataset for future training cycles. The validation of the framework is also applied to a real-world scenario, which is "Document Identification and Information Extraction". This problem originates mainly from banking systems, but can be extended anywhere. The method of classification and extraction of information has to be done correctly here. Experimental results show that including human feedback significantly enhances the ability of the model in order to increase robustness and accuracy in making decisions. The augmented approach, with a combination of machine efficiency and human insight, attains a higher learning standard-mainly in complex or ambiguous environments. The findings of this study show that human-in-the-loop reinforcement learning frameworks such as ARL can provide a scalable approach to improving model performance in data-driven applications.

强化学习人机协同决策优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。