← Writing

Hierarchical RL on OGBench

Skills, not steps

Flat reinforcement learning never solves OGBench's long manipulation tasks, even after 40 million steps. A library of small learned skills, stitched together by a learned composer, does. Here is how we got there and what broke along the way.

learned composer · success
Scene, task 5. Unlock the drawer, open it, put the cube inside, close and lock it, then unlock, open and re-lock the window. A learned composer picks every skill call; every skill is a learned policy. The side panel lists each call as it happens.

OGBench's scene environment gives a robot arm a desk with a cube, a drawer, a sliding window and two lock buttons. Task 1 is short: open the drawer and the window. Tasks 4 and 5 chain five to eight steps, and some of those steps make the desk look further from the goal before it gets closer. Those two tasks are where flat RL breaks, and where most of what follows is aimed.

The method is the same across every environment here, so it's worth laying out before any of the numbers.

A library of skills, and a composer that calls them

Instead of one policy for the whole task, we train a handful of short skills, each on its own with its own reward and its own check for when it's done. Then a second policy, the composer, learns which skill to call next. The composer never sees motor commands; its actions are skill calls.

press_button · 1.00 on its own
move_drawer · 1.00 on its own
move_window · 1.00 on its own
pick_place_cube · 0.95 on its own
place_in_drawer · 0.98 on its own · added mid-project

The loop that produced every result

Nothing here was hand-tuned into place. Every number in this post came out of the same four-stage loop: train each skill on its own, train a composer over skill calls, run an oracle plan through those same skills to find out which layer is broken, then intervene on whichever one the probe indicts.

Algorithm 1 · Decompose–compose–probe loop

Input: task M; initial skill vocabulary V

  1. for each k ∈ V: train πk in isolation
  2. repeat
  3. train composer μ over calls to V
  4. run oracle plan through V ⇒ ceiling c; failing step t*, skill k*, states S*
  5. if oracle succeeds(a)
  6. train μ longer, or fix its reward / executor
  7. else if πk* still improving(b)
  8. train πk* longer
  9. else if S* in-distribution for k*(c)
  10. train πk* further on S*
  11. else(d)
  12. specify new skill k'; train in isolation; V ← V ∪ {k'}
  13. until composed success plateaus

The probe on line 4 is what makes the rest work, so it's worth being exact about what the oracle is. The oracle supplies only the plan: the correct sequence of skill calls for the task — press_button(1), move_drawer(open), place_in_drawer(cube), and so on — the sequence a flawless composer would have chosen. It does not supply flawless control. Those calls are carried out by the same learned skill policies, through the same interface, as in any other episode. The oracle knows what to do next; the learned skills still have to do it, and they can still drop the cube.

That makes the oracle plan's score a ceiling — the best any composer could manage with the current library — and it separates two failures that look identical from outside. If the oracle plan succeeds, the skills can express the task and the fault is in the composer. If it fails, the skills cannot express it, and its failure trajectories point at the step, the skill called there, and the states it failed from. Branch (d) is the one that lets the method recover from a bad initial decomposition — a vocabulary that can't express the task shows up as an oracle failure with no capable skill, which is exactly the trigger for adding one.

The loop is run end to end by a single LLM agent (Claude Sonnet 5.5), with no human decisions inside it. The agent proposes the initial decomposition and each skill's specification, reads the probe's output to locate the failure, chooses the branch, and under branch (d) specifies the new skill. Humans set up the environments and each family's discretization; every decision after that was the agent's.

Two rules govern that decision. First, intervene on failure, not on distribution shift — a measurable shift between skill-training and deployment states is common, and is often not the cause. Second, never skip the probe: every time a layer was retrained without it, it was the wrong layer.

Flat RL learns the easy task, then stops

How much of that machinery is necessary? The baseline settles it. Standard PPO, given 40 million steps and OGBench's own reward, learns task 1 perfectly. On tasks 4 and 5 it never records a single success.

Flat PPO success on scene, over 40M environment steps

Training-time evaluation, 20 episodes per point. Tasks 4 and 5 sit on top of each other at 0.

A reasonable objection is that a hierarchy wins only because it makes fewer, coarser decisions. So we also ran flat PPO with each action held for 25 steps, which decides as often as our composer does. It stayed at 0.00 on task 4 across more than 40M steps. Coarse decisions are not what does the work.

Scene: all five tasks

With the library and a learned composer, the arm solves every scene task, including the two flat PPO never touches. Across all five tasks the best composer succeeds 86% of the time on held-out episodes.

task 1 · open
Task 1. Open the drawer and the window.
task 2 · unlock and lock
Task 2. Unlock both, close both, lock both again.
task 3 · rearrange
Task 3. Move the cube, close the window, open the drawer.
task 4 · put in drawer
Task 4. Unlock the drawer, open it, put the cube inside, close it.

Scene success per task

Held-out: 50 episodes per task on an unseen seed. Flat PPO was run on tasks 1, 4 and 5.

When the library is missing a skill

Our first library had no place_in_drawer, and tasks 4 and 5 sat at zero. Before retraining anything, we ran the probe from line 4: the oracle plan, executed through those same learned skills.

It only reached 0.28 — so the ceiling itself was the problem, and no amount of composer training could have reached past it. Digging in, pick_place_cube had put the cube into the drawer 0 times in 60. No amount of composer training could fix that, so we added a new skill. Its first two training runs scored zero, and the reason turned out to be the reference controller we used to measure it: it dropped the cube from tabletop height, which is too high for the drawer. Training at the depth the tasks actually use fixed it.

This is the part of the method we think matters most: the skill library isn't fixed up front. When the probe says no skill can do the job, the loop adds one, and the composer gets it as a new action.

The reward that paid the robot not to finish

The composer's first reward was shaped: a bonus for every part of the desk that matched the goal. That sounds helpful, but on tasks 4 and 5 the arm has to open the drawer, which is closed in the goal state. The shaping charged −1.05 for exactly that move. We repaired the potential, then tried removing shaping entirely and paying only for finishing.

Scene composer training, by reward

Training-time evaluation (8 episodes per task, moving average of 3). Held-out scores of each run's best checkpoint: binary 0.86, repaired 0.77, original 0.41.

Lights Out has the same trap. OGBench's reward counts lights that match the goal, but a correct solution often has to make the board look worse on the way. The learned composer's own solve goes from 5 wrong lights to 6 before it finishes. A shaped composer on the same board got stuck pressing one button 19 times in a row, flipping between 4 and 5 wrong lights.

Lights Out 4×4 composer training

Best seed of each, moving average of 3 evaluations. Final: binary 1.00, shaped 0.41.

Wrong lights after each press

The learned composer's 7-press solve of 4×4 task 5. Press 1 makes the board worse; a reward for matching lights would have punished it.

Lights Out: one skill, and where the composer runs out

OGBench's puzzle task is the classic Lights Out, played with a robot arm. Pressing a button flips it and its four neighbours; the arm has to reach a goal pattern. We need just one skill, press(k), which scores 1.00 on every board. On 3×3 and 4×4 the learned composer solves every task, and it also solves 4×4 boards it never trained on, 99% of the time.

learned composer
4×4, task 5. The learned composer, 7 presses.
unseen board
A 4×4 board it never trained on, also 7 presses.
exact planner + learned skill
4×6, 24 presses. With an exact planner — the oracle for this family — the same press skill solves all 20 tasks.

Learned composer success by board size

Best run per board, mean of its last 6 evaluations. The press skill and the oracle plan (here an exact planner over the board) both score 1.00 on every board, so the drop is the composer's.

Past 16 buttons the learned composer falls off a cliff. A solution is a set of buttons, so there are 2n candidates: 65,536 on 4×4 and about 16.8 million on 4×6. With a reward only for finishing, the composer has to stumble on a first solve, and on the big boards it never does. Swapping PPO for SAC with hindsight relabeling didn't move the cliff.

Cube: splitting the skill is what makes it trainable

OGBench's cube tasks ask the arm to move 2 to 8 cubes onto target spots, including swaps and stacks. One policy trained to pick and place a cube never learned (0.055). Split into grasp, move_to and release, each part reached 0.95 to 1.00.

2 cubes
Two cubes. Learned grasp and carry, scripted release.
3 cubes
Three cubes. A red flash marks a bumped neighbour; every cube still ends on its target.
4 cubes
Four cubes. Two bumps along the way; all four end on target.
8 cubes
Eight cubes, the most crowded board.

Multi-cube success, fix by fix

Scripted plan through the learned skills, 20 tasks × 5 episodes. None of these fixes retrained a skill.

The biggest jump came from a one-line difference between training and deployment. During training, the arm lifted clear before every new grasp. The deployed loop never did, so each grasp started low, sweeping across a crowded table it had never seen. Adding the same retreat took success from 0.24 to 0.72 without retraining anything. A learned composer over these skills still scores 0.00 on the six hardest cube tasks, though even the oracle plan only reaches 0.27 there — so on those tasks the skills are the binding constraint, not the composer.

What's still hard

Every number is the best run of its configuration. Scene numbers are held-out (50 episodes per task, unseen seed); composer numbers are the mean of a run's last six evaluations.