评测手机端小模型当电话秘书的应答能力,关注用户是否认可其处理方式。
CallScreenBench: Benchmarking Small Language Models as Phone Secretaries
- 构建电话秘书评测基准,聚焦文本对话决策层,不依赖外部工具或语音交互。
- 3个4比特量化模型在服务性、召回率和合理性上表现差异明显,大模型整体更优。
- 提出守护性诊断机制,帮助识别无需权限的代理场景,适合隐私敏感应用。
能在手机端运行的极小语言模型(4比特量化,0.6-4B参数)正日益具备代表用户执行任务的能力,使本地化任务自动化成为可能。其中一项关键任务是接听未知来电——电话秘书需在无合作呼叫方、可能面对恶意对方的情况下,自主决定如何应答。评估重点不是任务完成度,而是主人是否会认可其代理行为。本文提出CallScreenBench基准,包含五组基于主人认可度设计的自动通话与记录评估指标,每组均配有反向指标与不确定性估计。未设定全局排行榜或综合得分。通过三类模型家族的成对4比特检查点测试发现,大模型在服务性、召回率与合理性等多指标上表现更优;而筛选判别力排序不同。纯诈骗侧真阳性率(TPR)显示普遍警惕倾向,且当合法侧误报纳入时,成对区分度发生变化。脚本化退化代理暴露下限,如“挂断并回播”策略实现实体召回1.000。质量指标与守护性通道分开报告,避免单一通过/失败评分掩盖权衡。
原文摘要 · Abstract (English)
Language models small enough to run on a handset, quantized to a few bits, are increasingly capable of acting on their user's behalf -- which makes on-device task automation newly plausible. One such task is answering the phone. A phone secretary takes an unknown inbound call on its owner's behalf, and unlike the agents most benchmarks evaluate, it has no cooperative caller-assigned task to complete: the caller holds the goal and may be an adversary, while the secretary must begin deciding how to respond without an oracle. What matters is not task success but whether the owner would endorse how their proxy handled the call. We evaluate only the text-domain conversational decision layer; speech recognition, audio interaction, end-to-end latency, and handset execution are outside scope. We present CallScreenBench, which reports five automated call-and-note measure groups motivated by owner endorsement. Each is paired, where available, with a counter-metric and an uncertainty estimate; no benchmark-wide Q1-Q5 composite or leaderboard score is defined. Three guardedness diagnostics identify candidate cases for a toolless proxy that holds no credentials and calls no tools. Across three model families represented by paired 4-bit checkpoints (0.6-4B), the primary scoring snapshot gives the larger checkpoint higher point estimates on several service, recall, and plausibility measures, while triage discrimination follows a different ordering. Bare scam-side TPR rewards universal suspicion, and pairwise separation changes when legitimate-side false positives are included and across judge snapshots. Scripted degenerate agents expose further floors, including a hangup-and-echo policy with entity recall 1.000. We report quality measures and guardedness channels separately so that a single pass/fail score does not hide their trade-offs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。