Open field notes for reinforcement learning
Understand the world before you train the agent.
Practical, step-by-step environment specifications—observations, actions, rewards, edge cases and baselines—written to be built from.
Environment loop · live tracestable
Policy
π(a | s)
agent / step 0482
actionstate
World
LunarDock
reward +1.82
OBS 8DACT 2DCAP 1K
Field index · 00 guides