Fine-Tuning VLAs with Self-Demonstrated Generative Control for Multi-Task Manipulation
Distilling a pretrained VLA's priors through self-supervised online rollouts on the target embodiment yields a multi-task policy that fine-tuning on expert data alone cannot.
- Distills pretraining skills. Pick-and-place succeeds 0% → 55% with no expert demonstrations for those tasks — instruction following 0% → 90%.
- Boosts expert-taught skills. On the tasks that do have teleoperation (2.2 min of it), gear insertion goes 40% → 90%.
- Preserves held-out skills. On Push and Pour — never in any fine-tuning mixture — 6 blind raters score the requested motion attempted 38% → 82% (push) and 43% → 94% (pour).
Abstract
State-of-the-art vision-language-action (VLA) models such as π0.5 exhibit strong semantic understanding, instruction following and task behavior. However, when deployed on new robots, even minor mismatches in hardware configuration relative to pretraining can cause severe performance drops. Finetuning the VLA on in-domain expert data from the new embodiment improves performance on the expert task but leads to a loss in its original instruction following and behavioral priors. In this paper, we propose a self-supervised method that generates online interaction rollouts from the zero-shot VLA as additional training data for finetuning. Our experiments show this yields strong multi-task manipulation policies on the target robot that (1) inherit pretrained tasks distilled from the zero-shot VLA without expert demos for them, (2) improve success and generalist instruction following beyond expert-only fine-tuning on expert-teleoperated tasks, and (3) mitigate forgetting of held-out tasks from the pretrained policy, not part of fine-tuning. We demonstrate the success of our approach across test sets probing generalization on a real ALOHA robot and a new simulation benchmark in RoboTwin.
Method
On the real ALOHA robot we evaluate six task families: pick-up objects, pick-and-place objects, pick-and-place laundry, and bimanual colored gear insertion — plus push and pour / turn-over, two skill families held out of every fine-tuning mixture, expert and self-supervised alike, which probe what survives of the pretrained VLA's prior.
Results
Simulation Evaluation Protocol
Task suite used for our RoboTwin simulation benchmark, spanning self-supervised (TSS), expert-supervised (TES), novel-object (TNO), and novel-composition (TNC) tasks.
| Experiment | TSS (old) | TES (new) | TNO | TNC | Overall |
|---|---|---|---|---|---|
| π0.5 RoboTwin Base Policy mid-training, before post-training | 90.8 | — | — | — | — |
| Multi-Task ES+ES (Rehearsal Oracle) | 85.6 | 97.0 | 30.0 | 15.3 | 57.0 |
| Multi-Task ES+ES (Rehearsal + LoRA Oracle) | 84.6 | 77.0 | 29.5 | 11.3 | 50.6 |
| Multi-Task ES | 16.6 | 93.0 | 57.0 | 6.7 | 43.3 |
| Multi-Task ES+SS (Ours) | 70.6 | 98.0 | 44.5 | 14.0 | 56.8 |
Oracle baselines rehearse stored expert data from mid-training. Ours: rehearsing 10 self-demos per task recovers 82.5% of the rehearsal oracle’s TSS performance, landing 0.2 points below it Overall and above Rehearsal + LoRA. Average test set SR (%), 50 seeds per task.
What does a self-demonstration actually contain?
The base policy almost never succeeds on our robot — 0 successful grasps in 15 rollouts — yet its attempts are far from random. It reaches for the object the instruction names and moves toward the correct container; only the physical grasp fails. That is the signal self-demonstrations carry.
13 of 15 rollouts retry the task. Counts are after semantic and safety filtering.
Citation
@misc{garg2026finetuningvlasselfdemonstratedgenerative,
title={Fine-Tuning VLAs with Self-Demonstrated Generative Control for Multi-Task Manipulation},
author={Prachi Garg and Steve Xing and Prahit Yaugand and Saurabh Gupta and Derek Hoiem},
year={2026},
eprint={2608.19490},
archivePrefix={arXiv},
primaryClass={cs.RO},
url={https://arxiv.org/abs/2608.19490},
}