RoboCuber

A trained bimanual policy solving the Rubik’s cube

2026 · Source on GitHub

Loading…
0%
starting…

Drag the faces of the small cube to scramble it, then press Solve.

How it works

A classical solver (Kociemba’s algorithm) figures out which sequence of moves will solve the cube, and a neural network executes each one. The network sees the current arm positions, the cube’s pose, and which move to make, and it outputs joint targets for both arms. It’s a small MLP, about 14 million parameters, no vision and no pretraining. It looks at the last four frames of state and predicts a chunk of 40 actions ahead, executing the first one before replanning.

Training

I started with about 10,000 demonstrations from a scripted controller, a hand-coded routine that knows how to do each move but has no learning. I added noise to its actions so the policy would see what going slightly off-track looks like and learn to recover. Behavior cloning on this dataset, using action chunking, got the policy to where it could flip the cube between its hands but couldn’t pull off the harder turning moves. So I ran several rounds of DAgger, letting the policy go until it got stuck, then having the scripted expert take over and finish from that state. Those recovery trajectories went back into the training set. That pushed the success rate to about 45%, where it plateaued.

To get from there to 98%, I switched to PPO with GAE, running 1024 parallel sims on a single L4 GPU. The full run was about 150 million environment steps, roughly 26 GPU hours. Naively applying PPO would have destroyed the cloned policy almost immediately, so I took some care to preserve what it already knew. I warmed up the value function before letting gradients touch the policy, and added a KL penalty against the original cloned weights, similar to how RLHF fine-tunes language models.

The reward is mostly sparse. The policy gets a bonus when it completes a move, with light shaping in between. Small rewards for tipping the cube toward the right face and rotating the top layer toward the target angle keep the policy from ever being too far from a learning signal. The only big penalty is for knocking the cube out of the workspace.

That got individual moves to nearly 100%, but only 13% of full solves went through. During training, every move started with the arms in their home position, but in a real solve each move picks up wherever the last one left off. The policy had never practiced from those states. Retraining with episodes that started mid-solve fixed this, bringing full solves to 98%.

Code

The source is on GitHub.