arXiv:2601.06054cs.CLcs.AI2026-01综述

用推理大模型自动审查营销内容合规性,提升审核效率与准确性。

A Multi-Stage Workflow for the Review of Marketing Content with Reasoning Large Language Models

  • 分阶段流程利用微调推理模型识别文本合规问题。
  • 小模型先生成推理过程再输出判断,提升准确性。
  • 对比不同训练策略与奖励函数组合,优化审核效果。

推理型大语言模型在解决复杂任务上表现出色。本文提出并评估了一种多阶段工作流,利用微调后的推理大模型辅助营销内容审核,确保其符合指定要求。贡献包括:(i) 提出一种不依赖外部知识表示的自动识别文本合规问题的新方法;(ii) 比较监督微调(SFT)与组相对策略优化(GRPO)在该任务中的有效性;(iii) 评估小模型在给出最终回答前生成推理链条的效果;(iv) 分析不同奖励函数及其组合对GRPO训练模型性能的影响。

原文摘要 · Abstract (English)

Reasoning Large Language Models (LLMs) have shown promising results when tasked with solving complex problems. In this paper, we propose and evaluate a multi-stage workflow that leverages the capabilities of fine-tuned reasoning LLMs to assist in the review process of marketing content, making sure they comply with a given list of requirements. The contributions of this paper are the following: (i) we present a novel approach -- that does not rely on any external knowledge representation -- for the automatic identification of compliance issues in textual content; (ii) compare the effectiveness of different fine-tuning strategies like Supervised Fine-Tuning (SFT) and Group Relative Policy Optimization (GRPO) in training models to solve this problem; (iii) we evaluate the effectiveness of training small LLMs to generate reasoning tokens before providing their final response; (iv) we evaluate how the choice and combinations of different reward functions affects the performance of a model trained with GRPO.

内容审核推理模型营销合规LLM应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。