ACCRUEContinual Off-Policy Reinforcement Learning for Robot Foundation Models

1 KAIST2 NVIDIA

ACCRUE enables robot foundation models to learn new tasks faster through off-policy RL while retaining previously learned skills.

31.0–68.5%

less robot interactionAutoMate · new tasks vs. single-task RL

38.7–42.9%

less robot interactionNIST · new tasks vs. single-task RL

98.3%

mean final prior-task success118 / 120 trials · 6 prior tasks

Real-world learning

ACCRUE vs. single-task RL

π0.5 + FastTD3 · real robot

Interactive learning curves

Task learning curve

Single-task RL ACCRUE
Measured success against online robot interaction timeSelect a task to compare single-task RL and ACCRUE.
— min
Metric definitions

Success rate uses a rolling window of 20 valid online RL episodes and the existing EMA display smoothing (0.8 new + 0.2 previous). Dashed segments use the available episodes before the first complete window. Robot time includes online warmup, but not demonstration collection or wall-clock training. Shading marks the difference in total recorded robot interaction time. Stage 1 is shared, so it has no separate ACCRUE run. Smoothed endpoints can be below the unsmoothed 95% stopping criterion.

Download the plotted values (JSON)

Prior-task retention

Earlier tasks evaluated over 20 trials per checkpoint.

Earlier tasks evaluated after stage 5.

20 trials per task
View every retention checkpoint

The task just learned is not counted as a prior task. Blank cells indicate that a task was not yet eligible for retention evaluation.

Video overview

Supplementary video03:00 · English narration · CCTranscriptDownload video

Method

Actor–critic inheritance
ACCRUE foundation model and compact actor–critic architecture from the approved method presentation

Update the RFM with accumulated demonstrations and inherit the actor–critic.

Simulation experiments

FORGE with GR00T + FastSAC, evaluated separately from the real-world setup.

New-task learning

FORGE / FWT ↑

RL: 4 task orders × 4 seeds.

Metric and comparison scope

FWT is the mean peak-held current-task learning-curve AUC, not a success-rate delta. The demonstration-only RFM is a static reference from 12 task evaluations; RL uses four task orders × four seeds. This focused comparison does not claim the highest FWT among all continual methods: ER has FWT 0.833. These point estimates have no displayed uncertainty; see the paper for full results.

Qualitative prior-task rollouts after learning NutThread.

ACCRUE

Continual Off-Policy Reinforcement Learning for Robot Foundation Models

Read PDF

BibTeX

Expanded ACCRUE method diagram