首个评测波斯语音视频模型的基准,涵盖诗歌、音乐等文化任务。
PARSA-Bench: A Comprehensive Persian Audio-Language Model Benchmark
- 构建16项任务,覆盖语音理解、语调分析与波斯文化音频识别
- 现有模型在诗歌格律检测上接近随机猜测,说明对韵律感知能力不足
- 适合研究跨语言多模态、波斯语文化理解或语音模型评估的学者
波斯语因古典诗歌、传统音乐和普遍的语码转换而具有独特的语音理解挑战,现有基准均未涵盖这些特性。我们提出PARSA-Bench(波斯语音推理与语音评估基准),是首个针对波斯语语言与文化评估大音视频模型的基准,包含16个任务和超过8000个样本,覆盖语音理解、副语言分析及文化音频理解。其中10项为新引入任务,包括诗歌格律与风格识别、传统波斯音乐理解、语码转换检测。文本基线模型始终优于音频模型,表明当前模型可能无法有效利用转录之外的音频信息。文化相关任务暴露了明显的能力缺陷:所有模型在格律(vazn)检测上表现接近随机水平,无论规模大小,说明当前模型仍难以捕捉韵律感知。数据集已公开于 https://huggingface.co/datasets/MohammadJRanjbar/PARSA-Bench。
原文摘要 · Abstract (English)
Persian poses unique audio understanding challenges through its classical poetry, traditional music, and pervasive code-switching - none captured by existing benchmarks. We introduce PARSA-Bench (Persian Audio Reasoning and Speech Assessment Benchmark), the first benchmark for evaluating large audio-language models on Persian language and culture, comprising 16 tasks and over 8,000 samples across speech understanding, paralinguistic analysis, and cultural audio understanding. Ten tasks are newly introduced, including poetry meter and style detection, traditional Persian music understanding, and code-switching detection. Text-only baselines consistently outperform audio counterparts, suggesting models may not leverage audio-specific information beyond what transcription alone provides. Culturally-grounded tasks expose a qualitatively distinct failure mode: all models perform near random chance on vazn detection regardless of scale, suggesting prosodic perception remains beyond the reach of current models. The dataset is publicly available at https://huggingface.co/datasets/MohammadJRanjbar/PARSA-Bench
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。