用AI辅助标注稳定大模型,降低人工成本并提升可靠性。
AI-Powered Annotation Pipelines for Stabilizing Large Language Models: A Human-AI Synergy Approach
- 结合弱监督与置信度标注的AI管道,自动识别模型不稳定性模式。
- 在多个稳定性维度上实现持续校准,显著提升模型一致性与准确性。
- 适合医疗、金融等对可靠性要求高的行业应用。
大语言模型在高度监管行业中因不稳定、推理不一致、幻觉和性能波动等问题而难以部署,尤其在工作流中表现突出。现有稳定化方法如基于人类反馈的强化学习(RLHF)和监督微调虽有量化改进,但依赖大量人工标注,成本高且难持续扩展。本文提出一种基于AI的标注流水线,系统性识别、标记并修复大模型输出中的不稳定性模式。通过人机协同机制,结合自动化弱监督与置信度标注,并由目标人类验证确保反馈信息的可靠性和道德合规性。引入语义一致性、事实正确性与逻辑连贯性三类稳定性标注,构建反馈闭环,实现模型的持续校准与鲁棒性增强。
原文摘要 · Abstract (English)
LLM implementations are failing in highly regulated industries owing to instability issues, inconsistent reasoning, hallucinations and performance variability, especially in workflows. These reliability issues restrict safe use of LLM in areas that need the precision of facts and consistent behavior (Aiyappa et al., 2023). The current methods of stabilization, such as, reinforcement learning with human feedback (RLHF) and supervised fine-tuning, offer quantifiable improvements but are expensive and based on the intensive annotation of humans, thus being not easily scaled in a sustainable way (Dong et al., 2023; Retzlaff et al., 2024). This paper presents an AI-based annotation pipeline that systematically identifies, labels, and fixes for instability patterns on LLM output. Our human-AI synergy method combines the models of automated weak supervision and confidence-based annotation with the target human validation to guarantee the reliability and moral uprightness of feedback information (Cabitza et al., 2023; Jiang et al., 2023). The semantic consistency, factual correctness, and logical coherence categories of stability-specific annotation are introduced into our framework, allowing the continuous calibration of models and the enhancement of their robustness based on the feedback loops (Honovich et al., 2021; Nan et al., 2021).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。