RACECalibrated Cumulative Scoring
for Rollout-Free Checkpoint Selection

Changyeon Kim1Yekyung Nah1Kimin Lee1,2
1 KAIST2 Config

RACE ranks robot policy checkpoints using successful reference trajectories, without deploying every candidate.

0.83

rank correlation in simulationHeld-out demos · Spearman ρ · 95% CI [0.72, 0.92]

8 tasks × 10 seeds

simulation checkpoint ranking12 checkpoints per run · 3 reference sources

1,800

real-robot evaluation trials6 tasks · 2 policy families · 30 trials / checkpoint

Real-world success vs. RACE

Measured success and RACE scores across training checkpoints.

Put Two Cubes · π₀.₅-DROID00:20 · English narration · CCDownload video
PnP Three Fruit · π₀.₅-DROID00:20 · Timelapse · No audioDownload video

Illustrative task footage, not individual checkpoint recordings. Success and RACE use separate scales.

Explore the checkpoint measurements ↓

Real-world checkpoint ranking

How closely does an offline score track measured deployment success?

Franka Research 3 · real robot

Joint-velocity control · 10 checkpoints per task · 30 trials per checkpoint.

Illustrative task footage; checkpoint success rates come from separate evaluations.

Explore the measured checkpoints

Loading measured results…
Each point is one checkpoint

Metric definitions and evaluation protocol

The paper protocol uses 20 training demonstrations, 30 held-out successful reference demonstrations, and 10 candidate checkpoints per task, with 30 robot trials per checkpoint (1,800 total trials). The available plotting bundle contains 9 checkpoints per π₀.₅-DROID task and 10 per Diffusion Policy task; this page displays those 57 measured points without filling missing values. RACE uses task-specific scoring settings. The scatter plot uses within-task z-normalization; selecting a task shows its own ranking metrics. All tasks shows the mean of task-level metrics, not correlations recomputed from pooled points.

Spearman ρ measures rank agreement. MMRV (Mean Maximum Rank Violation) measures success-rate gaps caused by ranking inversions. nRegret measures the gap between the selected checkpoint and the best measured checkpoint, normalized by the task’s success-rate range. RACE scores are ranking signals, not success probabilities.

Download measured values and metrics (JSON)
Compare all four criteria
View checkpoint values

Video overview

The idea and experiments in three minutes.

Supplementary video03:00 · English narration · CCTranscriptDownload video

Lower error can still mean failure

One fatal mistake can hide behind a low average error. Larger, recoverable deviations can still succeed.

Counterexample: Policy A has lower average prediction error but fails after one fatal deviation; Policy B deviates more but succeeds.
View example at full size

How RACE scores a checkpoint

  1. Calibrate locally

    Set action-group thresholds from nearby successful reference states.

  2. Score deviations

    Treat within-threshold differences equally; penalize larger departures.

  3. Accumulate over time

    Early low scores reduce later contributions. Average to rank checkpoints.

RACE pipeline: calibrate action-group deviations, then aggregate timestep scores using a running product.

Simulation experiments

Checkpoint ranking across eight DexMimicGen tasks and three sources of successful references.

Mean Maximum Rank Violation ↓

Lower is better. Lines show 95% hierarchical bootstrap intervals over 10 seeds per task and 8 tasks.

Full simulation results and confidence intervals

Means are averaged over seeds within each task, then equally over eight tasks. RACE uses task-specific settings. Noised rollouts are successful trajectories collected under perturbations; this evaluates reference-source shift, not robustness to failed demonstrations. Hit@k is the fraction of runs in which the empirically best checkpoint is among the criterion’s top k candidates.

Eight manipulation tasks

Task illustrations from DexMimicGen, not RACE-evaluated rollouts. Source website: CC BY-SA 4.0. Clips resized and re-encoded.

What makes the score useful?

Temporal scoring and consequential deviations.

Accumulation improves ranking

Same timestep scores, cumulative aggregation: ρ 0.63 → 0.76; nRegret 0.180 → 0.096.

Four-task temporal ablation · paper Table III(b).

Similar MSE, different outcomes

Similar MSE changes, but success drops 4 vs. 40 points. RACE drops 1.42% vs. 52.69%.

Box Cleanup · uniform vs. bottleneck perturbations · 50 paired trials · Table II.

Scope and limitations

RACE provides relative checkpoint rankings, not calibrated probabilities of success. Variation in successful reference actions is a proxy for tolerable deviations; it does not establish recovery or predict behavior at unseen states. Selected checkpoints still require deployment validation. Main results use task-specific scoring settings; fully shared settings reduce mean simulation rank correlation from 0.80 to 0.69.

RACE

Calibrated Cumulative Scoring for Rollout-Free Checkpoint Selection

Read PDF

BibTeX