CallScreenBench: Benchmarking On-Device Models as Phone Secretaries
Researchers introduce CallScreenBench, a benchmark to evaluate on-device AI models for handling unknown phone calls autonomously.
- CallScreenBench is the first benchmark designed to evaluate on-device AI models for autonomous phone call handling.
- The benchmark focuses on the opening turn of a call, where the caller’s intent is ambiguous or adversarial.
- Unlike traditional benchmarks, success is measured by user endorsement rather than task completion.
- The work underscores the potential for privacy-preserving, on-device AI in real-world applications.
A new benchmark called CallScreenBench has been proposed to assess the capabilities of on-device language models as phone secretaries. These models, quantized to run efficiently on smartphones, must handle unknown inbound calls without prior context or user guidance. Unlike traditional AI benchmarks that focus on task completion, CallScreenBench evaluates whether the model’s response aligns with what the user would endorse, even in adversarial or uncooperative scenarios. The benchmark emphasizes the opening turn of the call, where the caller’s intent is often unclear, making it a unique challenge for on-device AI systems. The work highlights the growing feasibility of on-device task automation, particularly for privacy-sensitive applications like call screening, where cloud-based processing may not be desirable or possible.
Provides a standardized way to test and improve small language models for real-world, privacy-sensitive tasks like call screening.
Highlights opportunities for companies to deploy on-device AI solutions that reduce reliance on cloud processing and enhance user privacy.
Demonstrates the practical capabilities of tiny AI models in everyday scenarios.
- Quantized models
- AI models compressed to use fewer bits per parameter, enabling efficient execution on low-power devices like smartphones.
- On-device AI
- Artificial intelligence systems that run locally on a device rather than relying on cloud servers.
UT artificial intelligence researchers awarded grants in Department of Energy initiative - The Daily Texan
Artificial Intelligence And Growing Biosecurity Concerns – Analysis - Eurasia Review
From NASA to the Classroom: the Engineer Bringing AI to Those Left Behind - United Nations Sustainable Development Group
UEmbed: Unified Sparse and Dense Multimodal Embeddings
CoWAM: Coordination Contracts for Selective Policy Intervention with WAMs
FTC Inquiry into AI ‘Ideological Bias’ Draws First Amendment Objections - Broadband Breakfast
The US Federal Trade Commission (FTC) has launched an inquiry into AI 'ideological bias', prompting concerns from free speech advocates.
Austin leaders to get report on residents' priorities for AI governance - KEYE
Austin city leaders will receive a report on residents' top priorities for AI governance, aiming to shape the city's AI development.
SecurityDisrupting a Criminal Scam Operation
OpenAI shut down accounts linked to a Cambodia-based criminal network using ChatGPT for romance and investment scams.
UK's first class of students aiming for a bachelor's degree in artificial intelligence set to begin studies - WUKY
The UK's first class of students is set to begin studying for a bachelor's degree in artificial intelligence. This marks a significant step in the country's efforts to develop AI talent.
Artificial intelligence: Why firms are struggling to set prices - BBC
Companies are struggling to set prices due to artificial intelligence. Firms are finding it difficult to balance pricing strategies with AI-driven insights.
BusinessAfter killer quarter, Palantir CEO Alex Karp calls AI industry ‘Marxist’
Palantir’s CEO Alex Karp criticized AI frontier labs as untrustworthy despite the company’s $1 billion profit this quarter.