Forward-process reinforcement learning for robot manipulation

RoboFFT: Fine-tuning Generative Robot Policies with Forward Process

*Equal contribution   †Corresponding authors
1Institute of AI, Peking University   2PKU-PsiBot Joint Lab   3Nanjing University

RoboFFT overview

We present RoboFFT, an online reinforcement learning framework for finetuning generative robot policy with forward process. (a) Reverse finetuning formulates the reverse denoising process as a multi-step MDP, while forward finetuning aligns with the same objective as training by following the forward noising process. (b) A unified framework across popular generative robot policy on simulation benchmarks. (c) Real-world deployment with integration of real world RL. (d) Improved performance in success rate, training stability and efficiency.

The problem

Improving generative policies beyond demonstrations

Diffusion and flow-based policies capture complex, multimodal action distributions from demonstrations. Yet imitation-only policies remain limited by imperfect demonstrations and distribution shifts, and further improvement typically requires additional expert data.

Reinforcement learning offers a way to improve through environment interaction and reward feedback. For generative robot policies, however, likelihood estimation is generally intractable, making standard likelihood-based policy-gradient updates difficult to apply.

RoboFFT is an online reinforcement learning framework that applies forward noising to sampled actions and uses the weighted score/flow matching loss to construct a surrogate policy ratio for PPO-style updates. This reuses the generative pretraining objective for reinforcement fine-tuning. Experiments with popular generative robot policies on simulation benchmarks, including long-horizon and sparse-reward tasks, demonstrate improved success rates, training stability, and efficiency. We also evaluate VLA fine-tuning and integrate RoboFFT into a real-world RL framework for robot deployment.

4policy optimization settings
flow · diffusion · VLA simulation · real robot
1.53x / 38.0x
training efficiency gainvs. ReinFlow / state-task DPPO
45.0% → 93.8%
mean real-world success ratefour tasks, pretrained → online RoboFFT

Project overview

Introduction Video

The central idea

Method

RoboFFT pipeline

RoboFFT pretrains a generative policy, collects rewarded rollouts, applies forward noising to sampled actions, and uses the matching loss for PPO-style updates.

For a rollout action chunk $a_t^0$, the same forward path used in generative pretraining produces a noisy action:

$$a_t^k = \alpha_k a_t^0 + \sigma_k\epsilon, \qquad \epsilon\sim\mathcal{N}(0,I).$$

Prior work connects weighted denoising and flow-matching objectives to a variational evidence lower bound. We therefore use the matching loss as a surrogate negative log-likelihood for the clean action, rather than treating it as an ordinary auxiliary regression loss:

$$\mathcal{L}_\theta^w \approx -\operatorname{ELBO}_\theta(a_t^0\mid o_t)+c(a_t^0,o_t).$$

Taking the matching-loss difference between the current and rollout policies cancels the parameter-independent term and constructs a surrogate policy ratio:

$$\widehat{\rho}_t(\theta)=\exp\left(\mathcal{L}_{\theta_{\mathrm{old}}}^{w}(a_t^0,o_t)-\mathcal{L}_{\theta}^{w}(a_t^0,o_t)\right).$$

We estimate each matching loss using $N_{\mathrm{mc}}$ sampled noise levels and Gaussian perturbations per action, reusing the same samples for the current and old policies to reduce variance. A ratio temperature $\alpha_{\mathrm{ratio}}$ rescales this loss difference before exponentiation, giving $\widehat{\rho}^{\,s}_t$. This calibrated ratio is optimized with PPO clipping and the rollout advantage:

$$J_{\mathrm{RoboFFT}}=\mathbb{E}\left[\min\left(\widehat{\rho}^{\,s}_t\widehat{A}_t,\,\operatorname{clip}(\widehat{\rho}^{\,s}_t,1-\epsilon,1+\epsilon)\widehat{A}_t\right)\right].$$

This is the central distinction of RoboFFT: the policy-gradient ratio is built from differences in the same matching objective used for generative pretraining. RoboFFT-F applies this update to flow policies, while RoboFFT-D uses noise-prediction losses for diffusion policies.

Exploration and stabilization. RoboFFT-F injects Gaussian noise with strength $\eta$ into its flow sampler during rollouts (SDE sampling). RoboFFT-D weights denoising errors by $\widetilde{\Delta\lambda}$, the change in clipped log signal-to-noise ratio between adjacent diffusion steps, to account for the noise schedule. It also attenuates negative-advantage terms by $\beta_{\mathrm{neg}}\in(0,1]$; $\beta_{\mathrm{neg}}=1$ leaves them unchanged. On harder tasks, an anchor loss regularizes the policy toward actions from its best-evaluation checkpoint. The analyses below examine these design choices and how the resulting updates change action quality.

Reverse-process RL

Optimizes transitions through the internal denoising chain, producing a multi-step MDP and a heavier optimization graph.

RoboFFT forward-process RL

Uses the matching loss as a surrogate action negative log-likelihood, then forms the policy ratio from its current-versus-old loss difference.

Main evidence

Results

Robomimic: four sparse-reward tasks

Lift · Can · Square · Transport
RoboFFT state and pixel results on Robomimic

Across state and pixel observations, RoboFFT is competitive on easier tasks and shows more sustained improvement on the long-horizon Square and Transport tasks.

Training efficiency

Policy training time per iteration
RoboFFT training efficiency

Mean and standard deviation over the first ten training iterations.

RoboFFT-F achieves a mean 1.53x speedup over ReinFlow. RoboFFT-D achieves a mean 38.0x speedup over DPPO on Robomimic state-input tasks.

VLA fine-tuning

RoboFFT VLA fine-tuning results

VLA fine-tuning success-once results on LIBERO-Object and LIBERO-Spatial.

RoboFFT improves the pretrained $\pi_0$ policy through its flow-based action head and remains competitive with the $\pi_{RL}$ baseline.

Qualitative evaluation

Simulation Rollouts

Lift Grasp and lift
Can Pick and place
Square Precision assembly
Transport Long-horizon bimanual task

Deployment under limited interaction

Real-World Results

We integrate RoboFFT into a real-world RL framework, starting from a rectified-flow policy pretrained on thirty SpaceMouse demonstrations. IQL estimates advantages from collected rollouts, while offline RoboFFT, imitation refresh, and online refinement are repeated for five iterations with approximately 30 rollouts per iteration.

Demonstrations → Offline RoboFFT + IQL → Real-robot rollouts → IL refresh → Online refinement
TaskRobotTypePretrainOfflineOnline
Pick EggplantFranka Research 3Pick-and-place12/2018/2020/20
Peg InsertionFranka Emika PandaPrecision assembly10/2019/2020/20
PourFranka Emika PandaGranular manipulation9/2015/2018/20
Cabinet RetrievalFranka Research 3Long horizon5/2014/2017/20

Pick Eggplant: post-training comparison

PretrainBC AggregationDRWROffline RoboFFTOnline RoboFFT / RL100
12/2014/2012/2018/2020/20

All methods use the same data budget; this separates RL post-training from simple data aggregation.

Pick Eggplant
Peg Insertion
Pour
Cabinet Retrieval
Four real-world RoboFFT task rollouts

Successful real-world rollouts across the four-task suite.

Selected analysis

What drives the improvement?

The following Square-state ablations examine the exploration and matching-loss estimation choices introduced above, followed by the diffusion-specific weighting and negative-advantage stabilization.

SDE exploration noise

For RoboFFT-F, $\eta$ controls the Gaussian noise injected into the rollout sampler. Moderate noise balances exploration and stable policy improvement; too little or too much degrades performance.

SDE exploration noise ablation

Monte Carlo estimation

$N_{\mathrm{mc}}$ is the number of forward-noising samples used to estimate each action’s matching loss. For RoboFFT-F, increasing it from 1 to 2 improves convergence, while larger budgets add limited gains.

Monte Carlo repetition ablation

Log-SNR weighting

For RoboFFT-D, $\widetilde{\Delta\lambda}$ weights denoising errors across noise levels in the surrogate loss. It improves convergence over plain weighting; this controls loss estimation, not rollout exploration noise.

RoboFFT-D log-SNR weighting ablation

Negative-advantage discount

For RoboFFT-D, $\beta_{\mathrm{neg}}$ scales negative-advantage update terms to limit over-penalization. Compared with no discount ($\beta_{\mathrm{neg}}=1$), $\beta_{\mathrm{neg}}=0.125$ gives the strongest sustained performance in this Square-state ablation.

RoboFFT-D negative-advantage discount ablation
Surrogate ratio discrepancy analysis

Surrogate-ratio discrepancy

How closely does the matching-loss ratio used in Method agree with a likelihood-based reference? During RoboFFT-F finetuning on Square-state, we compare the temperature-scaled surrogate with a ReinFlow-style closed-form reference on 2,048 action chunks per iteration, using the same current and old policies.

Spearman correlation measures agreement in sample ranking; relative-error quantiles measure numerical differences between ratios. Near the end, relative errors fall below 5% for at least 90% of chunks, while rank correlation remains around 0.2. Numerical agreement improves in this setting, without establishing strict equivalence.

Action quality and diversity rollout timeline

Action value and diversity

To examine how the updates reshape the policy, we compare pretrained and finetuned action samples at the same states and with matched latent noises along a successful Square rollout. An auxiliary critic $Q_{\mathrm{viz}}(s,a)$ scores action quality for analysis only; $\Delta Q_{\mathrm{viz}}$ is the finetuned-minus-pretrained mean score.

Diversity $D$ is the mean pairwise squared distance between action chunks, normalized by their dimension; $\Delta D=D_{\mathrm{pre}}-D_{\mathrm{ft}}$ measures its reduction. Around grasping, transport, and alignment, higher diagnostic values and lower diversity suggest that finetuning concentrates samples around more useful behaviors.

Open research

Resources

Reproducibility

Code, checkpoints, configurations, benchmark instructions, and experiment logs will be linked here after public release.

BibTeX

@article{robofft,
  title={RoboFFT: Fine-tuning Generative Robot Policies with Forward Process},
  author={Li, Yu and Hu, Shenghe and Wang, Yuhan and Pu, Yaoxiang and Zhang, Haotong and Chen, Yuanpei and Yang, Yaodong},
  year={2026}
}