Reward Atlas
All field guidesEsc to close
Sign in
LLM / AgenticPublicPublishedv1.0Advanced

SQL-Detective-v1

A multi-turn agent answers analytics questions by querying a read-only SQLite sandbox. Tool use, verifiable rewards, 12-turn budget. Maintained by Bhaskar & Jaswanth · updated 28 Aug 2026.

Dataset
2,400 q · 12 DBs
Action
tool call | final answer
Turn cap
12
Best baseline
GRPO-7B 61.4%
09

Validation & testing

Grader tests

pytest tests/grader
  • Row order does not affect exact matchPASS
  • Duplicate rows are counted (multiset)PASS
  • Canary tokens zero the rewardPASS
  • Floating-point tolerance 1e-6WARN

    Revenue sums can differ at the 7th decimal between SQLite builds.

Sandbox isolation

Databases are mounted read-only and the container runs with networking disabled. ATTACH, PRAGMA writes and extension loading are rejected before execution.

Guide details
Version
Type
LLM / Agentic
API
OpenEnv
License
Apache-2.0
Seeds
3
Domains
tool-usedatacoding
Install
pip install reward-atlas-sql-detective

Issue with this step? Suggest an edit