Fine-Tuning VLAs with Self-Demonstrated Generative Control for Multi-Task Manipulation

Prachi Garg, Steve Xing*, Prahit Yaugand*, Saurabh Gupta, Derek Hoiem
University of Illinois Urbana-Champaign
*Equal contribution
TL;DR

Distilling a pretrained VLA's priors through self-supervised online rollouts on the target embodiment yields a multi-task policy that fine-tuning on expert data alone cannot.

  1. Distills pretraining skills. Pick-and-place succeeds 0% → 55% with no expert demonstrations for those tasks — instruction following 0% → 90%.
  2. Boosts expert-taught skills. On the tasks that do have teleoperation (2.2 min of it), gear insertion goes 40% → 90%.
  3. Preserves held-out skills. On Push and Pour — never in any fine-tuning mixture — 6 blind raters score the requested motion attempted 38% → 82% (push) and 43% → 94% (pour).

Overview

Abstract

State-of-the-art vision-language-action (VLA) models such as π0.5 exhibit strong semantic understanding, instruction following and task behavior. However, when deployed on new robots, even minor mismatches in hardware configuration relative to pretraining can cause severe performance drops. Finetuning the VLA on in-domain expert data from the new embodiment improves performance on the expert task but leads to a loss in its original instruction following and behavioral priors. In this paper, we propose a self-supervised method that generates online interaction rollouts from the zero-shot VLA as additional training data for finetuning. Our experiments show this yields strong multi-task manipulation policies on the target robot that (1) inherit pretrained tasks distilled from the zero-shot VLA without expert demos for them, (2) improve success and generalist instruction following beyond expert-only fine-tuning on expert-teleoperated tasks, and (3) mitigate forgetting of held-out tasks from the pretrained policy, not part of fine-tuning. We demonstrate the success of our approach across test sets probing generalization on a real ALOHA robot and a new simulation benchmark in RoboTwin.

Method

We fine-tune on a union of two data sources on the same target robot. Expert demonstrations 𝒟ES come from human teleoperation. Self-demonstrations 𝒟SS are generated by rolling out the frozen zero-shot π0.5 on the target embodiment and keeping what survives semantic and safety filtering — no human effort and no reward model.

Self-supervised generative rollouts: the frozen policy deploys action chunks on the target robot to produce self-demonstrations, which are filtered and mixed with human expert demonstrations to form the training data.
Both supervision streams are collected on the target robot and merged into one training mixture.

On the real ALOHA robot, we evaluate across five task families — pick-up objects, pick-up and place objects, pick-and-place laundry in basket, bimanual colored gear insertion, and push objects (a skill held out of every fine-tuning mixture, probing retention of the pretrained VLA's prior).

5
Task families
59
Unique task prompts
120
Test scenes

Results

Benchmark C — RoboTwin (Simulation)

Simulation Evaluation Protocol

Task suite used for our RoboTwin simulation benchmark, spanning self-supervised (TSS), expert-supervised (TES), novel-object (TNO), and novel-composition (TNC) tasks.

TSS Self-Supervised Tasks
place dual shoes
beat block hammer
pick dual bottles
place container plate
handover mic
place empty cup
place bread basket
lift pot
open microwave
put bottles dustbin
TES Expert Supervised Tasks
stack blocks red on green
stack blocks green on red
TNO Novel Object Tasks
stack cylinder on cube
stack sphere on cube
stack sphere on cube, distractor
stack cylinder on cube, distractor
TNC Novel Composition Tasks
stack blocks three
stack bowls two
stack bowls three

What does a self-demonstration actually contain?

The base policy almost never succeeds on our robot — 0 successful grasps in 15 rollouts — yet its attempts are far from random. It reaches for the object the instruction names and moves toward the correct container; only the physical grasp fails. That is the signal self-demonstrations carry.

Per-episode timeline of grasp and place attempts across 15 base-policy rollouts, marked by whether each attempt targeted the correct object and container.
Every grasp and place attempt across 15 base-policy rollouts, positioned by when it occurs in the episode. PDF
94%of grasp attempts target the requested object
89%of place attempts target the correct container
0successful grasps in 15 base-policy rollouts

13 of 15 rollouts retry the task. Counts are after semantic and safety filtering.

Citation

@misc{garg2026finetuningvlasselfdemonstratedgenerative,
      title={Fine-Tuning VLAs with Self-Demonstrated Generative Control for Multi-Task Manipulation},
      author={Prachi Garg and Steve Xing and Prahit Yaugand and Saurabh Gupta and Derek Hoiem},
      year={2026},
      eprint={2608.19490},
      archivePrefix={arXiv},
      primaryClass={cs.RO},
      url={https://arxiv.org/abs/2608.19490},
}