arXiv:2604.21345cs.AIcs.CL2026-04

构建可复用的会议摘要评估流水线,验证多模型表现差异。

Evaluating AI Meeting Summaries with a Reusable Cross-Domain Pipeline

  • 设计可复用评估系统,固定生成候选摘要并结构化标注真值。
  • 114场会议测试中,GPT-5.1在内容完整性和覆盖度上领先。
  • 适合关注AI评估流程标准化与跨领域复用的研究者。

工业团队常在缺乏稳定回归或模型选择评估的情况下部署大语言模型功能。本文提出一个可复用的AI会议摘要评估系统,包含结构化真值构建、固定候选生成、基于事实的评分、持久化报告以及隐私约束的在线监控与提名接口。在线证据不构成基准:隐私安全的聚合导出可实现活跃监控、严苛场景检测与方向性趋势追踪,而不暴露客户数据。离线路径在114场会议(city_council、private_data、whitehouse_press_briefings)上进行评估,生成340个会议-模型配对和680次GPT-4.1-mini、GPT-5-mini、GPT-5.1的评委打分。在固定协议下,各模型准确率差异不显著(校正p值0.053–0.448),但GPT-4.1-mini平均准确率最高(0.583);在保留率方面,GPT-5.1在完整性(0.886)和覆盖度(0.942)上显著领先。类型化分析指出whitehouse_press_briefings为高难度场景,后续对GPT-4.1、GPT-5-mini、GPT-5.4的重测复用相同流程与评委。本扩展预印本保持核心结果一致,补充了来自配套论文的跨领域复用细节及深度对比分析。

原文摘要 · Abstract (English)

Industrial teams often deploy large language model features before stable regression or model selection evaluation exists. We present a reusable evaluation system for AI meeting summaries that combines structured ground-truth (GT) construction, fixed candidate generation, claim-grounded scoring, persisted reporting, and a privacy-bounded online monitoring and nomination interface. The online evidence is not itself a benchmark: privacy-safe aggregate exports show active monitoring, hard regime detection, and directional movement without exposing customer data. We benchmark the offline path on 114 meetings across city_council, private_data, and whitehouse_press_briefings, yielding 340 completed meeting-model pairs and 680 judge runs for gpt-4.1-mini, gpt-5-mini, and gpt-5.1. Under this fixed protocol, accuracy differences are not statistically significant under Holm correction (corrected p-values 0.053-0.448), although gpt-4.1-mini has the highest mean accuracy (0.583); the significant separation is on retention, where gpt-5.1 leads on completeness (0.886) and coverage (0.942). Typed slices isolate whitehouse_press_briefings as an accuracy-hard regime, and a later focused rerun over gpt-4.1, gpt-5-mini, and gpt-5.4 reuses the same stack under the same judges and metrics. This extended preprint keeps those core results aligned with the formal submission while adding a more detailed repository-level account of cross-domain reuse from the companion AI-search paper and an additional typed DeepEval contrastive analysis. Model naming note. Running text uses canonical model names on first introduction. Tables, filenames, and artifact IDs retain compact report labels for consistency with the packaged benchmark outputs. Table A maps the two conventions and is repeated in Section 4.3 where candidate-generation settings are defined.

评估系统会议摘要模型评测可复用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。