用大模型当会议代表,测试其代参会能力。
MEETING DELEGATE: Benchmarking LLMs on Attending Meetings on Our Behalf
- 构建真实会议记录的基准测试,评估模型代参会表现。
- 约60%回复覆盖关键点,但存在重复与无关内容。
- 适合想减轻会议负担的团队或研究者参考。
在现代职场中,会议是思想交流和团队对齐的重要方式,但常面临耗时、调度冲突和参与低效等问题。大语言模型(LLMs)在自然语言生成与推理方面表现突出,引发疑问:能否让它们有效代为参会?为此,我们开发了一个基于LLM的会议代表原型系统,并利用真实会议转录文本构建了全面的基准测试。评估显示,GPT-4/4o在积极与谨慎策略间保持平衡;Gemini 1.5 Pro倾向保守,而Gemini 1.5 Flash和Llama3-8B/70B则更主动。总体上,约60%的回复至少覆盖一个关键点。然而,仍需改进以减少无关或重复内容,并提升对实际场景中常见转录错误的容忍度。此外,我们在真实环境中部署系统并收集演示反馈。结果表明,使用LLM作为会议代表具有潜力,但也面临挑战,为实际应用提供了重要洞见。
原文摘要 · Abstract (English)
In contemporary workplaces, meetings are essential for exchanging ideas and ensuring team alignment but often face challenges such as time consumption, scheduling conflicts, and inefficient participation. Recent advancements in Large Language Models (LLMs) have demonstrated their strong capabilities in natural language generation and reasoning, prompting the question: can LLMs effectively delegate participants in meetings? To explore this, we develop a prototype LLM-powered meeting delegate system and create a comprehensive benchmark using real meeting transcripts. Our evaluation reveals that GPT-4/4o maintain balanced performance between active and cautious engagement strategies. In contrast, Gemini 1.5 Pro tends to be more cautious, while Gemini 1.5 Flash and Llama3-8B/70B display more active tendencies. Overall, about 60\% of responses address at least one key point from the ground-truth. However, improvements are needed to reduce irrelevant or repetitive content and enhance tolerance for transcription errors commonly found in real-world settings. Additionally, we implement the system in practical settings and collect real-world feedback from demos. Our findings underscore the potential and challenges of utilizing LLMs as meeting delegates, offering valuable insights into their practical application for alleviating the burden of meetings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。