SimHum co-trains a bimanual manipulation policy on simulation data and human demonstrations, then fine-tunes it on a small real-robot dataset. Simulation supplies robot-valid actions. Human demonstrations supply real-world observations. This video walks through the recipe and the real-robot results. With only 80 real-robot episodes per task, SimHum reaches a 62.5% average success rate on held-out OOD scenes. That is 53.7% higher than Real only in absolute success rate. Under matched collection time, SimHum is 35.0% higher than the best single-source pre-training baseline in absolute success rate.
Abstract
Real-robot demonstrations are prohibitively expensive, while simulation data and real-world human demonstrations are both scalable but each leaves a distinct gap—simulation suffers from a sim-to-real visual gap (a), and human data suffers from a human-to-robot embodiment gap (b). In this work, we identify a natural yet underexplored complementarity between these sources: simulation contributes robot-valid actions absent in human data, while human data provides real-world observations that simulation struggles to render. Building on this insight, we present SimHum, a co-training recipe that extracts kinematic priors from simulation and visual priors from human observations, then fine-tunes on a small real-robot dataset (c). SimHum exhibits strong scene-generalizable and data-efficient capabilities. With only 80 real-robot episodes per task, it achieves 62.5% success on held-out OOD scenes across four bimanual tabletop tasks, 53.7% higher than Real only in absolute success rate. Moreover, in a controlled data-collection study with matched collection time, SimHum improves over the best single-source pre-training baseline by 35.0% in absolute success rate (d).
Data Collection
Three Complementary Data Sources
Real-robot data is the bottleneck. It is the only source that is both robot-valid and real, and it is the one that costs the most to collect. Simulation supplies robot-valid actions, but its rendering leaves a sim-to-real visual gap. Human demonstrations supply real-world observations, but a human hand leaves a human-to-robot embodiment gap. The two gaps do not overlap, so the two cheap sources cover each other.
We run three collection pipelines. Simulation episodes are generated automatically on top of RoboTwin 2.0. Real-robot episodes are teleoperated on COBOT Magic through a leader–follower system. Human episodes are captured with a VR interface and a stationary camera. Simulation shares identical kinematics with the real robot. Human data shares an identical camera with the real robot. All three cover the same task set.
Hover the connectors to explore how the sources relate
Simulation
Real Robot
Human
Collected Data Visualization
Explore our multi-source demonstration dataset captured across diverse environments. Switch between the tabs below to view episodes for different manipulation tasks.
Simulation Data
Human Data
Real Robot Data
Approach
We employ a Two-Stage Training Paradigm to train our Modular Policy Architecture. In the Sim-and-Human Pre-training stage (a), we leverage Modular Action Encoders/Decoders to extract transferable kinematic priors from simulation, and Domain-specific Vision Adaptors to extract visual priors from human data. Subsequently, for Real-robot Fine-tuning (b), we restructure the policy by selectively retaining the compatible components—specifically the Real-world Vision Adaptor and Robot Encoder/Decoder—to achieve data-efficient generalization.
SimHum in Real-World Scenarios
We evaluate the SimHum under In-Distribution (ID) and Out-of-Distribution (OOD) settings. While baselines degrade significantly in OOD scenarios (featuring unseen background textures, distractors, and extreme lighting), SimHum maintains robust performance. With only 80 real-robot episodes per task, SimHum reaches a 62.5% average OOD success rate. That is 53.7% higher than Real only in absolute success rate.
All Videos Autonomous 1×
In-Distribution (ID)
Out-of-Distribution (OOD)
Decoupling the Effects: Why Both Sources Matter
Hover or tap the left chart to read exact values (mean ± standard error)
Data Efficiency and Performance Scalability
Hover or tap the charts to read exact values (mean ± standard error)
All three studies are run on Stack Bowls Two under OOD.
Baseline Failure Cases in OOD Settings
We visualize typical failure cases of baseline methods in Out-of-Distribution (OOD) scenarios. Please switch between the tabs below to view different manipulation tasks.
Real-only
Trained only on limited real-world data, it overfits to spurious visual correlations and generalizes poorly.
HumReal
Pre-trained on human data and then fine-tuned on limited real-world data, it struggles with the embodiment gap due to kinematic mismatches between human and robots.
SimReal
Pre-trained on simulation and then fine-tuned on limited real-world data, it is hindered by the visual gap, as simulated rendering does not perfectly align with the real world.
All Videos Autonomous 1×
Real Only
HumReal
SimReal
BibTeX
@inproceedings{fang2026simhum,
title = {Sim-and-Human Co-training for Data-Efficient and Scene-Generalizable Bimanual Manipulation},
author = {Fang, Kaipeng and Liang, Weiqing and Li, Yuyang and Zhang, Ji and Zeng, Pengpeng and
Shen, Heng Tao and Song, Jingkuan and Gao, Lianli},
booktitle = {10th Annual Conference on Robot Learning},
year = {2026},
}
Human Data Collection Setup
The system has three parts: hardware, a real-time GUI, and synchronized recording. A Meta Quest 3 tracks the hand, a stationary ego camera matches the robot's head view, and a foot pedal marks episode boundaries. GUI telemetry and the scene view are recorded frame-synchronized.