用分段评估+主动学习,让大模型自检长文本更准更省
Monocle: Hybrid Local-Global In-Context Evaluation for Long-Text Generation with Uncertainty-Based Active Learning
- 分段评估+全局整合,避免长文本评价失真
- 引入人类反馈提升评分与人工判断一致率
- 基于不确定性选样本,标注成本降低30%以上
评估长文本生成质量极具挑战性,尤其在输入长度增加时,现有基于大模型评分的方法性能会下降。为此,我们提出一种分而治之的评估策略:将整体评价任务拆分为多个局部评分任务,再进行全局整合。该方法使每个文本片段独立评估连贯性与质量,同时兼顾全文结构一致性。此外,我们设计了一种混合式上下文学习方法,利用人类标注增强局部与全局评价性能,使模型评分更贴近人类判断。最后,开发了基于不确定性的主动学习算法,智能筛选需人工标注的样本,在实际应用中显著降低标注成本。实验表明,该框架优于多个代表性基线方法,验证了其有效性。
原文摘要 · Abstract (English)
Assessing the quality of long-form, model-generated text is challenging, even with advanced LLM-as-a-Judge methods, due to performance degradation as input length increases. To address this issue, we propose a divide-and-conquer approach, which breaks down the comprehensive evaluation task into a series of localized scoring tasks, followed by a final global assessment. This strategy allows for more granular and manageable evaluations, ensuring that each segment of the text is assessed in isolation for both coherence and quality, while also accounting for the overall structure and consistency of the entire piece. Moreover, we introduce a hybrid in-context learning approach that leverages human annotations to enhance the performance of both local and global evaluations. By incorporating human-generated feedback directly into the evaluation process, this method allows the model to better align with human judgment. Finally, we develop an uncertainty-based active learning algorithm that efficiently selects data samples for human annotation, thereby reducing annotation costs in practical scenarios. Experimental results show that the proposed evaluation framework outperforms several representative baselines, highlighting the effectiveness of our approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。