Counterfactual Video Generation Enables
Scalable Humanoid Loco-Manipulation

1 Amazon FAR2 UC Berkeley3 Stanford University4 Carnegie Mellon University† Amazon FAR team co-leads.

CoRL 2026Conference on Robot Learning (CoRL 2026)

tl;dr: We expand 4 real video demonstrations into diverse training data for humanoid loco-manipulation.

We train only on videos of boxes and generated videos of bins, barrels, and balls.

Real-world results

Our policy handles these real objects without real-world fine-tuning.

  • Onboard depth onlyNo MOCAP or reference motion
  • A single autonomous policyWith joystick steering.
  • Trained from video aloneZero-shot Sim2Real

Explore objects, scale, strategy, pose, and robustness.

Teaching humanoids to pick up diverse everyday objects

Humans generalize from rich interactions. How can robots gain this experience?

Visual imitation gives robots a starting point: VideoMimic learns climbing and sitting from human video.

Scaling this approach means collecting more demonstrations. Common solutions are:

(1) Collect your own data: motion capture and recording are costly and do not scale.

(2) Internet video: abundant, but reconstruction-friendly demonstrations are hard to find.

Our solution: Expand a few recordings into new interactions with a video model.

See different applications

Human-scene interaction

Vary the stairs

Human-object interaction

Vary the object

Counterfactual video generation

The video model

Video-to-video generation.

What PRISM does

We sample interactions, then run real-to-sim.

Original recordingCounterfactual videos
Original video A person picks up a box
Generated videoA different shape
Generated videoA different grasp
Generated videoA different reach

Counterfactual videos

These videos show interactions that did not happen in the source recording, but could have happened with a different object.

Prompt template

<VIDEO>

In a real-world continuous footage. Preserve the reference video's original background, lighting and camera viewpoint.

replace the box with <CLS>, pick it up and carry with two hands.

<VIDEO> = source recording · <CLS> = box, bin, barrel, or ball

Why video-to-video

A source video supplies the person, scene, and action. Text-to-video must create all three.

Method

We reconstruct these interactions in 3D, retarget them to the robot, and train in simulation.

We correct generation and reconstruction errors for physically feasible motion and hand–object contact.

Generated CF video

3D reconstruction

Retargeted motion

Learn in simulation

Transfer to the real world

PRISM overview: monocular human–object videos are reconstructed in a shared world frame with contact anchors, retargeted to the humanoid, and used to train co-tracking teachers that are distilled into one depth-based student policy for zero-shot sim-to-real deployment.

See our paper for more details.

Watch the learned policy on the real robot ↑

Robustness and Generalization

Initial object-pose coverage

We compare object poses in 137 generated clips and 4 upright seed videos. The plots show positions relative to G1 and yaw counts in six 30° bins modulo 180°.

Initial object poses relative to G1: generated clips form a wider spread around the four upright seed poses, including lying-down objects. The orientation histogram totals 137 clips across six yaw bins.

Zero-shot elevated pick-up

We train only on flat terrain. In simulation, the policy succeeds zero-shot on a 35° ramp (top) and a 0.43 m support (bottom), but fails at 45° and 0.45 m.

Tall training objects may explain this transfer by exposing the policy to similar grasp heights. We can generate elevated pick-up demonstrations with simple V2V prompts; training on them may extend this range.

Two simulation sequences show G1 approaching, grasping, and lifting an object. Top: a yellow bin on a 35-degree ramp. Bottom: a yellow box on a 0.43-meter support. Arrows annotate contact forces.

Why video in the agentic era?

GPT-6 Astra (Extra High) for the same motion: a person picks up a bin, carries it, and sets it down.

Text (without video). We caption the generated video, then give Astra only that text to script the SMPL motion.

Video. A video model turns a real box-carrying demonstration into a bin-carrying video. Astra calls a Real2Sim module, then refines the reconstruction.

The simplest way is still to show a demonstration. Video still helps—and it should! The agentic pipeline only helps scale Real2Sim.

BibTeX

If you find our work helpful, please cite:

@article{wang2026counterfactual,
  title={Counterfactual Video Generation Enables Scalable Humanoid Loco-Manipulation},
  author={Wang, Zihan and Wu, Zhen and Abbeel, Pieter and Duan, Rocky and Malik, Jitendra and Sferrazza, Carmelo and Liu, C Karen and Shi, Guanya and Kanazawa, Angjoo},
  journal={arXiv preprint arXiv:2609.38172},
  year={2026}
}

Acknowledgements

We thank Chung Min Kim, Arthur Allshire, Hongsuk Choi, Isabella Yu, Junyi Zhang, Jacob Berg, Yen-Jen Wang, Sirui Chen, Charlie Cheng, Jiashun Wang, Siheng Zhao, Youjian Huang, JC Hu, Haochen Wang, Haozhi Qi and Qitao Zhao for their support and valuable feedback.