VisCodeBench: a new benchmark dataset for Text-to-Vis reflecting real-world practice
Abstract
Generating visualisations from natural language requests (text-to-vis) enables users to explore data more efficiently. While large language models (LLMs) can complete simple visualisation requests, their performance on challenging ones is not well covered by existing benchmarks. We introduce VisCodeBench, an end-to-end text-to-vis benchmark with 2,672 samples from three different domains: educational, scientific, and general audience scenarios with varying levels of complexity. We benchmark two representative coding agents, Codex-GPT5 and Codex-OSS20B, using six evaluation metrics that cover execution performance, chart content, and visual quality. Our results show that Codex-GPT5 is more robust in completing visualisation tasks. Codex-OSS20B exhibits lower execution reliability and weaker performance on complex workflows. Both are able to choose the right chart type, but have significant chart content errors, and both models are also slow, taking 50-90 seconds to produce results. The evaluation code, representative samples, and prompts are provided in the repository attached to this submission, while the full benchmark dataset is kept private and accessed via a leaderboard to prevent scraping and LLM training use.