Reward Atlas
All field guidesEsc to close
Sign in
LLM / AgenticPublicPublishedv1.0Advanced

SQL-Detective-v1

A multi-turn agent answers analytics questions by querying a read-only SQLite sandbox. Tool use, verifiable rewards, 12-turn budget. Maintained by Bhaskar & Jaswanth · updated 28 Aug 2026.

Dataset
2,400 q · 12 DBs
Action
tool call | final answer
Turn cap
12
Best baseline
GRPO-7B 61.4%
06

Starting state & episode end

Starting state

reset() samples a question uniformly from the split and mounts a fresh copy of its database, so no state leaks between episodes.

Episode boundaries

Terminated

The policy calls submit_answer. This is the only way to receive a non-zero reward.

Truncated

The 12-turn budget is exhausted. For GRPO-style training the reward is simply 0; for value-based methods, treat as truncation.

Guide details
Version
Type
LLM / Agentic
API
OpenEnv
License
Apache-2.0
Seeds
3
Domains
tool-usedatacoding
Install
pip install reward-atlas-sql-detective

Issue with this step? Suggest an edit