ConsumerBench: Benchmarking Generative AI Applications on End-User Devices
Abstract
The recent shift in Generative AI (GenAI) applications from cloud-only environments to end-user devices introduces new challenges in resource management and system efficiency. This paper presents ConsumerBench, a comprehensive benchmarking framework designed to evaluate the system efficiency and response time of GenAI models running on end-user devices. Unlike existing benchmarks that assume exclusive model access on dedicated GPUs, ConsumerBench simulates realistic multi-application scenarios executing concurrently on constrained hardware. Furthermore, ConsumerBench supports customizable workflows that simulate complex tasks such as multi-agent pipelines involving multiple GenAI models and cross-application dependencies. ConsumerBench monitors Service-Level Objectives (SLOs) of applications, as well as system-wide metrics like power, CPU/GPU utilization and memory bandwidth. Through extensive experiments, ConsumerBench reveals inefficiencies in existing GPU sharing strategies under realistic concurrent executions. The paper also provides practical insights for system and kernel designers, highlighting the importance of implementing SLO-aware GPU sharing strategies, and concurrency-aware kernel implementations to avoid resource contention.