Fine-tune-ready (0/1/5-shot) version of LLMSQL 1.0. LLMSQL 2.0 is a test-only benchmark and has no training data.
AI & ML interests
None defined yet.
Recent Activity
Organization Card
LLMSQL: Text-to-SQL benchmarks for the LLM era
GitHub · Documentation · PyPI · Kaggle
LLMSQL turns WikiSQL into benchmarks that are reliable for evaluating modern LLMs. The datasets come with an evaluation
environment: a pip-installable package (pip install llmsql) for inference and evaluation, documentation, and a task
for the LM Evaluation Harness.
Datasets
| Dataset | What it is | Use it for |
|---|---|---|
| llmsql-2.0 | LLMSQL 2.0. 2,000 hard, verified test questions over 942 Wikipedia tables: lookups, data conventions (e.g. winner-first scores), dates stored as text. Zero-shot, execution accuracy. | Evaluation (recommended) |
| llmsql-benchmark | LLMSQL 1.0. About 80k cleaned WikiSQL questions with train/val/test splits. | Training, and LLMSQL 1.0 results |
| llmsql-benchmark-finetune-ready | LLMSQL 1.0 as 0/1/5-shot prompt–completion pairs. | Fine-tuning |
| benchmark-evaluation-results | Model outputs behind the earlier LLMSQL leaderboard runs. | Reproducing published runs |
Quick start
from llmsql import inference_vllm, evaluate
inference_vllm("Qwen/Qwen2.5-Coder-7B-Instruct", output_file="outputs.jsonl", version="2.0")
print(evaluate("outputs.jsonl", version="2.0")["accuracy"])
Citation
@inproceedings{llmsql_bench,
title={LLMSQL: Upgrading WikiSQL for the LLM Era of Text-to-SQL},
author={Pihulski, Dzmitry and Charchut, Karol and Novogrodskaia, Viktoria and Koco{\'n}, Jan},
booktitle={2025 IEEE International Conference on Data Mining Workshops (ICDMW)},
year={2025},
organization={IEEE}
}
models 0
None public yet