arXiv:2412.18377cs.CLcs.AI2024-12NAACL被引 4

构建聊天机器人交互补全评估基准,助力智能输入优化

ChaI-TeA: A Benchmark for Evaluating Autocompletion of Interactions with LLM-based Chatbots

  • 提出聊天交互补全任务定义与评估框架
  • 在9个模型上测试,发现建议排序仍有提升空间
  • 适合对话系统、输入法研究者参考

大语言模型的兴起使越来越多的人机交互转向基于LLM的聊天机器人。这些模型支持用户以长篇多样化的自然语言表达各种话题和风格,但撰写消息耗时费力,亟需自动补全方案。本文提出聊天机器人交互补全任务,并构建ChaI-TeA:一个针对该任务的自动补全评估框架,包含任务明确定义、配套数据集与评价指标。我们用该框架评测了9个现成模型,结果表明当前模型表现尚可,但在生成建议的排序方面仍有显著提升空间。研究为从业者提供实践洞见,也为该领域开启新的研究方向。框架已开源,可作为后续研究基础。

原文摘要 · Abstract (English)

The rise of LLMs has deflected a growing portion of human-computer interactions towards LLM-based chatbots. The remarkable abilities of these models allow users to interact using long, diverse natural language text covering a wide range of topics and styles. Phrasing these messages is a time and effort consuming task, calling for an autocomplete solution to assist users. We introduce the task of chatbot interaction autocomplete. We present ChaI-TeA: CHat InTEraction Autocomplete; An autcomplete evaluation framework for LLM-based chatbot interactions. The framework includes a formal definition of the task, coupled with suitable datasets and metrics. We use the framework to evaluate After formally defining the task along with suitable datasets and metrics, we test 9 models on the defined auto completion task, finding that while current off-the-shelf models perform fairly, there is still much room for improvement, mainly in ranking of the generated suggestions. We provide insights for practitioners working on this task and open new research directions for researchers in the field. We release our framework to serve as a foundation for future research.

对话系统自动补全评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。