AnchorQ: Sample-Efficient Reinforcement Learning via Flow-Reversed Action Anchors

Anonymous Authors

Anonymous submission

How does AnchorQ work?

Real-world results

We evaluate three manipulation tasks on a Franka robot with a DROID setup, giving each RL method 100 online training episodes per task. Across 20 held-out configurations per task, AnchorQ outperforms frozen π0.5 and DSRL on all three tasks and is the only method to improve towel hanging over the base policy. Aggregate error bars show 95% Wilson confidence intervals.

Average error bars show 95% Wilson confidence intervals. p-values compare AnchorQ with each baseline using two-sided t-tests.

Showing AnchorQ on Stack Cups.

Simulation Results

We compare π0.5, DSRL, and AnchorQ with adaptive particle tilting on two LIBERO-90 task suites. Each plot shows individual task success across 100 episodes followed by the average across tasks, with mean ±1 standard error over four training seeds.

Success rate (%)
  • π0.5 zero-shot
  • DSRL
  • AnchorQ

Per-task success

LIBERO-90 task ID · hover or select a task for details

Average

All 15 tasks

Task and average error bars show ±1 SEM across four seeds. The average includes every task in the selected suite.

Watch the policies learn

500k

For AnchorQ rollouts, the colored overlays identify the nearest anchor and its reference direction from the executed latent.