AI & ML interests

None defined yet.

Recent Activity

pihull  updated a dataset about 24 hours ago
llmsql-bench/llmsql-2.0
pihull  updated a Space 1 day ago
llmsql-bench/README
pihull  updated a dataset 1 day ago
llmsql-bench/llmsql-2.0
View all activity

Organization Card

LLMSQL: Text-to-SQL benchmarks for the LLM era

GitHub · Documentation · PyPI · Kaggle

LLMSQL turns WikiSQL into benchmarks that are reliable for evaluating modern LLMs. The datasets come with an evaluation environment: a pip-installable package (pip install llmsql) for inference and evaluation, documentation, and a task for the LM Evaluation Harness.

Datasets

Dataset What it is Use it for
llmsql-2.0 LLMSQL 2.0. 2,000 hard, verified test questions over 942 Wikipedia tables: lookups, data conventions (e.g. winner-first scores), dates stored as text. Zero-shot, execution accuracy. Evaluation (recommended)
llmsql-benchmark LLMSQL 1.0. About 80k cleaned WikiSQL questions with train/val/test splits. Training, and LLMSQL 1.0 results
llmsql-benchmark-finetune-ready LLMSQL 1.0 as 0/1/5-shot prompt–completion pairs. Fine-tuning
benchmark-evaluation-results Model outputs behind the earlier LLMSQL leaderboard runs. Reproducing published runs

Quick start

from llmsql import inference_vllm, evaluate

inference_vllm("Qwen/Qwen2.5-Coder-7B-Instruct", output_file="outputs.jsonl", version="2.0")
print(evaluate("outputs.jsonl", version="2.0")["accuracy"])

Citation

@inproceedings{llmsql_bench,
  title={LLMSQL: Upgrading WikiSQL for the LLM Era of Text-to-SQL},
  author={Pihulski, Dzmitry and Charchut, Karol and Novogrodskaia, Viktoria and Koco{\'n}, Jan},
  booktitle={2025 IEEE International Conference on Data Mining Workshops (ICDMW)},
  year={2025},
  organization={IEEE}
}

models 0

None public yet