Long-WAM YAM

A YAM bimanual robot pretraining model trained on ABC 130K, with a LongLive 2.0 Robot-S video backbone. It predicts actions after imagining future video latents.

Download

hf download Efficient-Large-Model/Long-WAM-YAM --local-dir ./weights/Long-WAM-YAM

Inference

Load model.pt with the matching Long-WAM runtime, config.yaml and dataset_stats.json. Use 4 video denoising steps to sigma 0.9, followed by 10 action denoising steps.

The three RGB views use a fixed 384 × 320 layout:

+---------------------------------+
|       Top view: 256 × 320        |
+----------------+----------------+
| Left wrist     | Right wrist    |
| 128 × 160      | 128 × 160      |
+----------------+----------------+

P48 covers 1.6 seconds at 30 Hz. Sample history every four control steps: 13 images become four clean latent frames, followed by two imagined latents. Predict a 32-step action chunk, approximately 1.07 seconds.

State and action are 14-D: left arm's six joints and gripper, then right arm's six joints and gripper. Use the supplied z-score statistics to normalize state and denormalize predicted actions; RGB inputs use [-1, 1].

The matching runtime and base-model VAE/text assets are required separately. This is a pretraining checkpoint: task-specific adaptation and controlled closed-loop validation are required before real-robot use.

Downloads last month
15
Video Preview
loading

Dataset used to train Efficient-Large-Model/Long-WAM-YAM