用大模型自动评估日志摘要质量,无需参考文本
REFLEX: Reference-Free Evaluation of Log Summarization via Large Language Model Judgment
- 让大模型零样本判断摘要的相关性、信息量和连贯性
- 在多个数据集上表现稳定,比传统指标更精细
- 适合缺乏参考摘要的现实场景使用
日志摘要评估因缺乏高质量参考摘要,且现有指标如ROUGE、BLEU依赖表面词汇重合而受限。我们提出REFLEX,一种基于大语言模型(LLM)的无参考评估方法。REFLEX利用LLM作为零样本评价器,从相关性、信息量和连贯性等维度评估摘要质量,无需黄金标准参考或人工标注。实验表明,REFLEX在多个日志摘要数据集上能产生稳定、可解释且细粒度的评估结果,相比传统指标更有效地区分不同模型输出。该方法为现实中参考数据稀缺或不可用的场景提供了可扩展的评估方案。
原文摘要 · Abstract (English)
Evaluating log summarization systems is challenging due to the lack of high-quality reference summaries and the limitations of existing metrics like ROUGE and BLEU, which depend on surface-level lexical overlap. We introduce REFLEX, a reference-free evaluation metric for log summarization based on large language model (LLM) judgment. REFLEX uses LLMs as zero-shot evaluators to assess summary quality along dimensions such as relevance, informativeness, and coherence, without requiring gold-standard references or human annotations. We show that REFLEX produces stable, interpretable, and fine-grained evaluations across multiple log summarization dataset, and more effectively distinguishes model outputs than traditional metrics. REFLEX provides a scalable alternative for evaluating log summaries in real-world settings where reference data is scarce or unavailable.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。