Sim-and-Human Co-training for Data-Efficient and Scene-Generalizable Bimanual Manipulation

SimHum co-trains a bimanual manipulation policy on simulation data and human demonstrations, then fine-tunes it on a small real-robot dataset. Simulation supplies robot-valid actions. Human demonstrations supply real-world observations. This video walks through the recipe and the real-robot results. With only 80 real-robot episodes per task, SimHum reaches a 62.5% average success rate on held-out OOD scenes. That is 53.7% higher than Real only in absolute success rate. Under matched collection time, SimHum is 35.0% higher than the best single-source pre-training baseline in absolute success rate.

Abstract

Overview of SimHum: (a) the sim-to-real visual gap, (b) the human-to-robot embodiment gap, (c) the SimHum co-training framework, and (d) OOD generalization and data-efficiency results.

Real-robot demonstrations are prohibitively expensive, while simulation data and real-world human demonstrations are both scalable but each leaves a distinct gap—simulation suffers from a sim-to-real visual gap (a), and human data suffers from a human-to-robot embodiment gap (b). In this work, we identify a natural yet underexplored complementarity between these sources: simulation contributes robot-valid actions absent in human data, while human data provides real-world observations that simulation struggles to render. Building on this insight, we present SimHum, a co-training recipe that extracts kinematic priors from simulation and visual priors from human observations, then fine-tunes on a small real-robot dataset (c). SimHum exhibits strong scene-generalizable and data-efficient capabilities. With only 80 real-robot episodes per task, it achieves 62.5% success on held-out OOD scenes across four bimanual tabletop tasks, 53.7% higher than Real only in absolute success rate. Moreover, in a controlled data-collection study with matched collection time, SimHum improves over the best single-source pre-training baseline by 35.0% in absolute success rate (d).

Data Collection

Three Complementary Data Sources

Real-robot data is the bottleneck. It is the only source that is both robot-valid and real, and it is the one that costs the most to collect. Simulation supplies robot-valid actions, but its rendering leaves a sim-to-real visual gap. Human demonstrations supply real-world observations, but a human hand leaves a human-to-robot embodiment gap. The two gaps do not overlap, so the two cheap sources cover each other.

We run three collection pipelines. Simulation episodes are generated automatically on top of RoboTwin 2.0. Real-robot episodes are teleoperated on COBOT Magic through a leader–follower system. Human episodes are captured with a VR interface and a stationary camera. Simulation shares identical kinematics with the real robot. Human data shares an identical camera with the real robot. All three cover the same task set.

Hover the connectors to explore how the sources relate

Simulation pipeline: RoboTwin-OD assets, robot parameters, grasp pose generation and motion planning, domain randomization, simulation environment, trajectory collection.
Real-robot pipeline: COBOT Magic leader-follower teleoperation with a robot camera, leader arms and follower arms.
Human pipeline: a VR interface with a Meta Quest 3, a stationary ego camera and a recording pedal.

Simulation

Robot Action
Visual – sim2real visual gap
Low Cost – 2000 episodes in total
Identical Kinematics

Real Robot

Robot Action
Real-world Observation
High Cost – 320 episodes in total
Identical Camera

Human

Human Hand – embodiment gap
Real-world Observation
Low Cost – 2000 episodes in total

Collected Data Visualization

Explore our multi-source demonstration dataset captured across diverse environments. Switch between the tabs below to view episodes for different manipulation tasks.

Simulation Data

Human Data

Real Robot Data


Approach

SimHum architecture: (a) Sim-and-Human pre-training with modular action encoders and decoders plus domain-specific vision adaptors, and (b) real-robot fine-tuning with the retained modules.

We employ a Two-Stage Training Paradigm to train our Modular Policy Architecture. In the Sim-and-Human Pre-training stage (a), we leverage Modular Action Encoders/Decoders to extract transferable kinematic priors from simulation, and Domain-specific Vision Adaptors to extract visual priors from human data. Subsequently, for Real-robot Fine-tuning (b), we restructure the policy by selectively retaining the compatible components—specifically the Real-world Vision Adaptor and Robot Encoder/Decoder—to achieve data-efficient generalization.

SimHum in Real-World Scenarios

We evaluate the SimHum under In-Distribution (ID) and Out-of-Distribution (OOD) settings. While baselines degrade significantly in OOD scenarios (featuring unseen background textures, distractors, and extreme lighting), SimHum maintains robust performance. With only 80 real-robot episodes per task, SimHum reaches a 62.5% average OOD success rate. That is 53.7% higher than Real only in absolute success rate.

All Videos Autonomous 1×

In-Distribution (ID)

Base Scene
Complex Scene 1
Complex Scene 2
Complex Scene 3

Out-of-Distribution (OOD)

Extreme Complex Scene 1
Extreme Complex Scene 2

Decoupling the Effects: Why Both Sources Matter

Hover or tap the left chart to read exact values (mean ± standard error)

(a) Leave-one-factor-out ablation on human data. (b) Per-position progress rate on a 4x4 grid in an OOD scene for HumReal and SimHum.
Background diversity (F_bg)
SimHum (Full): 80.0% ± 8.9
w/o Factor: 20.0% ± 8.9
Distractor diversity (F_dis)
SimHum (Full): 75.0% ± 9.7
w/o Factor: 35.0% ± 10.7
Lighting diversity (F_light)
SimHum (Full): 75.0% ± 9.7
w/o Factor: 50.0% ± 11.2
Object diversity (F_obj)
SimHum (Full): 80.0% ± 8.9
w/o Factor: 50.0% ± 11.2

Data Efficiency and Performance Scalability

Hover or tap the charts to read exact values (mean ± standard error)

All three studies are run on Stack Bowls Two under OOD.

(a) Task success rate under matched 2, 4 and 8 hour collection-time budgets. (b) Task progress rate as real-robot demonstrations scale from 8 to 160. (c) Task progress rate as one pre-training source scales from 0 to 500 episodes.
2 h of collection time
Real only: 15.0% ± 8.0
SimHum (Ours): 30.0% ± 10.3
4 h of collection time
Real only: 15.0% ± 8.0
SimHum (Ours): 40.0% ± 11.0
8 h of collection time
HumReal: 20.0% ± 8.9
Real only: 25.0% ± 9.7
SimReal: 35.0% ± 10.7
SimHum (Ours): 70.0% ± 10.3
8 real demos
Real only: 38.3% ± 10.9
SimHum (Ours): 55.0% ± 11.1
19 real demos
Real only: 43.3% ± 11.1
SimHum (Ours): 60.0% ± 11.0
40 real demos
Real only: 45.0% ± 11.1
SimHum (Ours): 65.0% ± 10.7
80 real demos
Real only: 48.3% ± 11.2
SimHum (Ours): 83.3% ± 8.3
160 real demos
Real only: 58.3% ± 11.0
SimHum (Ours): 91.7% ± 6.2
0 pre-training episodes
Scaling Sim: 58.3% ± 11.0
Scaling Human: 60.0% ± 11.0
100 pre-training episodes
Scaling Sim: 61.7% ± 10.9
Scaling Human: 63.3% ± 10.8
200 pre-training episodes
Scaling Sim: 63.3% ± 10.8
Scaling Human: 70.0% ± 10.3
400 pre-training episodes
Scaling Sim: 75.0% ± 9.7
Scaling Human: 76.7% ± 9.5
500 pre-training episodes
Scaling Sim: 83.3% ± 8.3
Scaling Human: 83.3% ± 8.3

Baseline Failure Cases in OOD Settings

We visualize typical failure cases of baseline methods in Out-of-Distribution (OOD) scenarios. Please switch between the tabs below to view different manipulation tasks.

Real-only

Trained only on limited real-world data, it overfits to spurious visual correlations and generalizes poorly.

HumReal

Pre-trained on human data and then fine-tuned on limited real-world data, it struggles with the embodiment gap due to kinematic mismatches between human and robots.

SimReal

Pre-trained on simulation and then fine-tuned on limited real-world data, it is hindered by the visual gap, as simulated rendering does not perfectly align with the real world.

All Videos Autonomous 1×

Real Only

Final Score: 0/2

HumReal

Final Score: 1/2

SimReal

Final Score: 1/2

BibTeX

@inproceedings{fang2026simhum,
  title     = {Sim-and-Human Co-training for Data-Efficient and Scene-Generalizable Bimanual Manipulation},
  author    = {Fang, Kaipeng and Liang, Weiqing and Li, Yuyang and Zhang, Ji and Zeng, Pengpeng and
               Shen, Heng Tao and Song, Jingkuan and Gao, Lianli},
  booktitle = {10th Annual Conference on Robot Learning},
  year      = {2026},
}