测试大模型能否从嘈杂语音转录中提取语法正确句子
Investigating large language models for their competence in extracting grammatically sound sentences from transcribed noisy utterances
- 用波兰语数据集测试大模型从噪音对话中提取结构化语句的能力
- 部分提取结果语法错误,表明模型未完全掌握语法规则或无法有效应用
- 适合关注大模型语言理解局限性的研究者阅读
人类能轻松区分语音中的语义内容与填充停顿、不流畅和重述等噪声,这得益于对语法规则的内化。我们基于语言学实验,探究大语言模型(LLMs)是否具备类似能力,即从嘈杂对话的转录文本中提取语法正确的语句。在波兰语场景下开展两项评估实验,使用可能未被模型接触过的数据集以避免数据污染。结果显示,并非所有提取语句都语法正确,说明LLMs要么未充分掌握语法规则,要么虽掌握但无法有效运用。结论是:当前大模型对嘈杂语音的理解仍远不及人类水平。
原文摘要 · Abstract (English)
Selectively processing noisy utterances while effectively disregarding speech-specific elements poses no considerable challenge for humans, as they exhibit remarkable cognitive abilities to separate semantically significant content from speech-specific noise (i.e. filled pauses, disfluencies, and restarts). These abilities may be driven by mechanisms based on acquired grammatical rules that compose abstract syntactic-semantic structures within utterances. Segments without syntactic and semantic significance are consistently disregarded in these structures. The structures, in tandem with lexis, likely underpin language comprehension and thus facilitate effective communication. In our study, grounded in linguistically motivated experiments, we investigate whether large language models (LLMs) can effectively perform analogical speech comprehension tasks. In particular, we examine the ability of LLMs to extract well-structured utterances from transcriptions of noisy dialogues. We conduct two evaluation experiments in the Polish language scenario, using a~dataset presumably unfamiliar to LLMs to mitigate the risk of data contamination. Our results show that not all extracted utterances are correctly structured, indicating that either LLMs do not fully acquire syntactic-semantic rules or they acquire them but cannot apply them effectively. We conclude that the ability of LLMs to comprehend noisy utterances is still relatively superficial compared to human proficiency in processing them.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。