arXiv:2509.21732cs.CL2025-09EMNLP被引 2

测试大模型在对话转录上多问题问答的准确性,发现小模型可超越大模型。

How Accurate Are LLMs at Multi-Question Answering on Conversational Transcripts?

  • 对比多种大模型在相同对话上下文上回答多个问题的能力
  • 80亿参数微调公开模型准确率超过GPT-4o
  • 为低成本透明部署提供可行方案,适合工业应用

将大语言模型(LLMs)应用于长文本的多问题问答是重要挑战,尤其在工业场景中,高计算成本和延迟限制了其应用。本文系统评估了多种专有与开源模型在基于同一对话上下文回答多个问题任务上的表现。实验结果表明,尽管强效专有模型如GPT-4o整体表现最佳,但经过微调的开源模型(最大80亿参数)在准确率上已超越GPT-4o,展现出在真实场景中实现透明、经济高效的部署潜力。

原文摘要 · Abstract (English)

Deploying Large Language Models (LLMs) for question answering (QA) over lengthy contexts is a significant challenge. In industrial settings, this process is often hindered by high computational costs and latency, especially when multiple questions must be answered based on the same context. In this work, we explore the capabilities of LLMs to answer multiple questions based on the same conversational context. We conduct extensive experiments and benchmark a range of both proprietary and public models on this challenging task. Our findings highlight that while strong proprietary LLMs like GPT-4o achieve the best overall performance, fine-tuned public LLMs with up to 8 billion parameters can surpass GPT-4o in accuracy, which demonstrates their potential for transparent and cost-effective deployment in real-world applications.

大模型问答系统对话理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。