Hierarchical RL on OGBench
Skills, not steps
Flat reinforcement learning never solves OGBench's long manipulation tasks, even after 40 million steps. A library of small learned skills, stitched together by a learned composer, does. Here is how we got there and what broke along the way.
- 0.86success across all five scene tasks with our skill library
- 0.00flat PPO on the two hardest scene tasks, after 40M steps
- 1.00Lights Out 4×4, solved by a learned composer
OGBench's scene environment gives a robot arm a desk with a cube, a drawer, a sliding window and two lock buttons. Task 1 is short: open the drawer and the window. Tasks 4 and 5 chain five to eight steps, and some of those steps make the desk look further from the goal before it gets closer. Those two tasks are where flat RL breaks, and where most of what follows is aimed.
The method is the same across every environment here, so it's worth laying out before any of the numbers.
A library of skills, and a composer that calls them
Instead of one policy for the whole task, we train a handful of short skills, each on its own with its own reward and its own check for when it's done. Then a second policy, the composer, learns which skill to call next. The composer never sees motor commands; its actions are skill calls.
The loop that produced every result
Nothing here was hand-tuned into place. Every number in this post came out of the same four-stage loop: train each skill on its own, train a composer over skill calls, run an oracle plan through those same skills to find out which layer is broken, then intervene on whichever one the probe indicts.
Algorithm 1 · Decompose–compose–probe loop
Input: task M; initial skill vocabulary V
- for each k ∈ V: train πk in isolation
- repeat
- train composer μ over calls to V
- run oracle plan through V ⇒ ceiling c; failing step t*, skill k*, states S*
- if oracle succeeds(a)
- train μ longer, or fix its reward / executor
- else if πk* still improving(b)
- train πk* longer
- else if S* in-distribution for k*(c)
- train πk* further on S*
- else(d)
- specify new skill k'; train in isolation; V ← V ∪ {k'}
- until composed success plateaus
The probe on line 4 is what makes the rest work, so it's worth being
exact about what the oracle is. The oracle supplies only the
plan: the correct sequence of skill calls for the task —
press_button(1), move_drawer(open),
place_in_drawer(cube), and so on — the sequence a
flawless composer would have chosen. It does not supply flawless
control. Those calls are carried out by the same learned skill policies,
through the same interface, as in any other episode. The oracle knows
what to do next; the learned skills still have to do it, and they
can still drop the cube.
That makes the oracle plan's score a ceiling — the best any composer could manage with the current library — and it separates two failures that look identical from outside. If the oracle plan succeeds, the skills can express the task and the fault is in the composer. If it fails, the skills cannot express it, and its failure trajectories point at the step, the skill called there, and the states it failed from. Branch (d) is the one that lets the method recover from a bad initial decomposition — a vocabulary that can't express the task shows up as an oracle failure with no capable skill, which is exactly the trigger for adding one.
The loop is run end to end by a single LLM agent (Claude Sonnet 5.5), with no human decisions inside it. The agent proposes the initial decomposition and each skill's specification, reads the probe's output to locate the failure, chooses the branch, and under branch (d) specifies the new skill. Humans set up the environments and each family's discretization; every decision after that was the agent's.
Two rules govern that decision. First, intervene on failure, not on distribution shift — a measurable shift between skill-training and deployment states is common, and is often not the cause. Second, never skip the probe: every time a layer was retrained without it, it was the wrong layer.
Flat RL learns the easy task, then stops
How much of that machinery is necessary? The baseline settles it. Standard PPO, given 40 million steps and OGBench's own reward, learns task 1 perfectly. On tasks 4 and 5 it never records a single success.
Flat PPO success on scene, over 40M environment steps
Training-time evaluation, 20 episodes per point. Tasks 4 and 5 sit on top of each other at 0.
A reasonable objection is that a hierarchy wins only because it makes fewer, coarser decisions. So we also ran flat PPO with each action held for 25 steps, which decides as often as our composer does. It stayed at 0.00 on task 4 across more than 40M steps. Coarse decisions are not what does the work.
Scene: all five tasks
With the library and a learned composer, the arm solves every scene task, including the two flat PPO never touches. Across all five tasks the best composer succeeds 86% of the time on held-out episodes.
Scene success per task
Held-out: 50 episodes per task on an unseen seed. Flat PPO was run on tasks 1, 4 and 5.
When the library is missing a skill
Our first library had no place_in_drawer, and tasks 4 and 5
sat at zero. Before retraining anything, we ran the probe from line 4:
the oracle plan, executed through those same learned skills.
It only reached 0.28 — so the ceiling itself was the problem, and no
amount of composer training could have reached past it. Digging in,
pick_place_cube had put the cube into the drawer 0 times in
60. No amount of composer training could fix that, so we added a new
skill. Its first two training runs scored zero, and the reason turned out
to be the reference controller we used to measure it: it dropped the cube
from tabletop height, which is too high for the drawer. Training at the
depth the tasks actually use fixed it.
- 0.28oracle plan through the original skills — the ceiling
- 0 / 60cubes the old skill got into the drawer
- 0.25 → 0.75in-drawer placement after adding the new skill (p = 0.004)
This is the part of the method we think matters most: the skill library isn't fixed up front. When the probe says no skill can do the job, the loop adds one, and the composer gets it as a new action.
The reward that paid the robot not to finish
The composer's first reward was shaped: a bonus for every part of the desk that matched the goal. That sounds helpful, but on tasks 4 and 5 the arm has to open the drawer, which is closed in the goal state. The shaping charged −1.05 for exactly that move. We repaired the potential, then tried removing shaping entirely and paying only for finishing.
Scene composer training, by reward
Training-time evaluation (8 episodes per task, moving average of 3). Held-out scores of each run's best checkpoint: binary 0.86, repaired 0.77, original 0.41.
Lights Out has the same trap. OGBench's reward counts lights that match the goal, but a correct solution often has to make the board look worse on the way. The learned composer's own solve goes from 5 wrong lights to 6 before it finishes. A shaped composer on the same board got stuck pressing one button 19 times in a row, flipping between 4 and 5 wrong lights.
Lights Out 4×4 composer training
Best seed of each, moving average of 3 evaluations. Final: binary 1.00, shaped 0.41.
Wrong lights after each press
The learned composer's 7-press solve of 4×4 task 5. Press 1 makes the board worse; a reward for matching lights would have punished it.
Lights Out: one skill, and where the composer runs out
OGBench's puzzle task is the classic Lights Out, played with a robot arm.
Pressing a button flips it and its four neighbours; the arm has to reach a
goal pattern. We need just one skill, press(k), which scores
1.00 on every board. On 3×3 and 4×4 the learned composer solves
every task, and it also solves 4×4 boards it never trained on, 99% of
the time.
Learned composer success by board size
Best run per board, mean of its last 6 evaluations. The press skill and the oracle plan (here an exact planner over the board) both score 1.00 on every board, so the drop is the composer's.
Past 16 buttons the learned composer falls off a cliff. A solution is a set of buttons, so there are 2n candidates: 65,536 on 4×4 and about 16.8 million on 4×6. With a reward only for finishing, the composer has to stumble on a first solve, and on the big boards it never does. Swapping PPO for SAC with hindsight relabeling didn't move the cliff.
Cube: splitting the skill is what makes it trainable
OGBench's cube tasks ask the arm to move 2 to 8 cubes onto target spots,
including swaps and stacks. One policy trained to pick and place a cube
never learned (0.055). Split into grasp,
move_to and release, each part reached 0.95 to
1.00.
Multi-cube success, fix by fix
Scripted plan through the learned skills, 20 tasks × 5 episodes. None of these fixes retrained a skill.
The biggest jump came from a one-line difference between training and deployment. During training, the arm lifted clear before every new grasp. The deployed loop never did, so each grasp started low, sweeping across a crowded table it had never seen. Adding the same retreat took success from 0.24 to 0.72 without retraining anything. A learned composer over these skills still scores 0.00 on the six hardest cube tasks, though even the oracle plan only reaches 0.27 there — so on those tasks the skills are the binding constraint, not the composer.
What's still hard
- The learned composer for cube is far below the oracle plan through the same skills.
- The Lights Out composer fails beyond 16 buttons; bigger boards need better exploration, not a better skill.
- Cube's release step is scripted.
- Most skills were trained with one or two seeds; the scene composer with three.
Every number is the best run of its configuration. Scene numbers are held-out (50 episodes per task, unseen seed); composer numbers are the mean of a run's last six evaluations.