arXiv:2601.03531cs.CL2026-01被引 1

首个个性化音视频模型评测基准,专测模型理解个人语境能力。

PALM-Bench: A Comprehensive Benchmark for Personalized Audio-Language Models

  • 定义个性化音视频模型任务,聚焦个人概念识别与上下文推理。
  • 实测主流模型在跨任务个性化知识迁移上表现有限。
  • 适合研究个性化多说话人交互、具身智能的团队使用。

大型音视频模型(LALMs)在音频理解与生成方面表现强劲,但我们的全面评测发现,其行为仍以通用处理为主(如总结口语内容),难以有效支持个性化问答(如总结我好友所说的话)。人类则会基于个人背景来理解与决策。为此,我们正式提出个性化音视频模型(PALM)任务,旨在识别个人概念并进行个人上下文推理。同时,我们构建了首个基准测试(PALM-Bench),推动该领域方法进展,并在多说话人场景下对多个任务实现结构化评估。我们在代表性开源LALMs上进行了广泛实验,结果表明现有免训练提示与监督微调策略虽有提升,但在建模个性化知识及跨任务稳健迁移方面仍存局限。数据与代码将公开发布。

原文摘要 · Abstract (English)

Large Audio-Language Models (LALMs) have demonstrated strong performance in audio understanding and generation. Yet, our extensive benchmarking reveals that their behavior is largely generic (e.g., summarizing spoken content) and fails to adequately support personalized question answering (e.g., summarizing what my best friend says). In contrast, human conditions their interpretation and decision-making on each individual's personal context. To bridge this gap, we formalize the task of Personalized LALMs (PALM) for recognizing personal concepts and reasoning within personal context. Moreover, we create the first benchmark (PALM-Bench) to foster the methodological advances in PALM and enable structured evaluation on several tasks across multi-speaker scenarios. Our extensive experiments on representative open-source LALMs, show that existing training-free prompting and supervised fine-tuning strategies, while yield improvements, remains limited in modeling personalized knowledge and transferring them across tasks robustly. Data and code will be released.

音视频模型个性化评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。