06
Starting state & episode end
Starting state
reset() samples a question uniformly from the split and mounts a fresh copy of its database, so no state leaks between episodes.
Episode boundaries
Terminated
The policy calls submit_answer. This is the only way to receive a non-zero reward.
Truncated
The 12-turn budget is exhausted. For GRPO-style training the reward is simply 0; for value-based methods, treat as truncation.