arXiv:2503.01045cs.CL2025-03被引 4

用大模型自动评估多语言语音回忆,更贴近真实对话场景。

Language-agnostic, automated assessment of listeners' speech recall using large language models

  • 用大模型生成多语言故事并自动评分
  • 在嘈杂环境下仍能准确捕捉回忆顺序效应
  • 适合跨语言听力研究与临床评估使用

老年人普遍存在言语理解困难,传统测试因内容单一、仅限主流语言(如英语)而难以准确反映真实情况。本研究利用大型语言模型,在10种非英语母语者及英语母语者中,自动化生成高质量口语化故事,并对听后自由回忆进行高效评分。参与者在安静和背景噪音条件下,用母语听取并复述短篇故事。通过大模型文本嵌入与提示工程结合语义相似度分析,系统成功捕捉到时间顺序、首尾效应及噪声影响等已知认知规律,且不同语言间的回忆评分高度一致。该方法突破了传统测试对简单材料和封闭群体的依赖,可高精度映射不同语言、长度与细节的回忆数据,实现从故事生成到评分的全流程自动化,为自然语境下的言语理解评估提供了临床可行的解决方案。

原文摘要 · Abstract (English)

Speech-comprehension difficulties are common among older people. Standard speech tests do not fully capture such difficulties because the tests poorly resemble the context-rich, story-like nature of ongoing conversation and are typically available only in a country's dominant/official language (e.g., English), leading to inaccurate scores for native speakers of other languages. Assessments for naturalistic, story speech in multiple languages require accurate, time-efficient scoring. The current research leverages modern large language models (LLMs) in native English speakers and native speakers of 10 other languages to automate the generation of high-quality, spoken stories and scoring of speech recall in different languages. Participants listened to and freely recalled short stories (in quiet/clear and in babble noise) in their native language. LLM text-embeddings and LLM prompt engineering with semantic similarity analyses to score speech recall revealed sensitivity to known effects of temporal order, primacy/recency, and background noise, and high similarity of recall scores across languages. The work overcomes limitations associated with simple speech materials and testing of closed native-speaker groups because recall data of varying length and details can be mapped across languages with high accuracy. The full automation of speech generation and recall scoring provides an important step towards comprehension assessments of naturalistic speech with clinical applicability.

语音理解多语言大模型应用自动评分

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。