StatAI Lab

StatAI Lab and CMFAI Jointly Launch DataSciEval — A Unified Benchmark for LLM and AI Agent Data Science Capabilities

We are excited to announce the release of DataSciEval, a unified benchmark for evaluating the data science capabilities of large language models (LLMs) and AI agents, jointly developed by Shanghai University of Finance and Economics Stat-AI Lab (led by Professor Fan Zhou) and Hong Kong Polytechnic University CMFAI (led by Professor Jian Huang).

DataSciEval comprises 1,900 theoretical and methodological test tasks, along with 641 application tasks based on real-world datasets, covering foundational knowledge, methodological understanding, research-level derivation, code execution, data analysis, and report generation.

Unlike traditional benchmarks that evaluate models solely based on final answers, DataSciEval simultaneously incorporates theoretical reasoning processes, code execution processes, and final output quality into its evaluation framework, providing a more comprehensive characterization of model capabilities in real-world data science workflows.

Learn more at https://huggingface.co/spaces/StatAILab/DataSciEval.

Next post
Two Papers Accepted to STAl-X 2026 — One Selected for Paper Award