Diffusion Policy for Push-T

Re-implemented Diffusion Policy for the Push-T benchmark with a Conditional U-Net 1D, cosine noise schedule, EMA, and DDIM sampling, preserving multimodal action distributions instead of collapsing to a mean.

Details

Re-implemented Diffusion Policy (Chi et al., RSS 2023) for the Push-T benchmark: instead of predicting actions directly, the model learns to denoise a noisy action horizon conditioned on recent observations, which makes it well suited to multimodal behavior where the same state may admit several valid actions. Architecture: Conditional U-Net 1D with down dimensions [256, 512, 1024], kernel size 5, and FiLM conditioning. Global conditioning is built from the first n_obs_steps observations. Training: epsilon prediction with MSE loss, cosine noise schedule via diffusers.DDPMScheduler, EMA (inv_gamma=1.0, power=0.75), and AdamW (lr=1e-4, betas=(0.95, 0.999), weight_decay=1e-6). Horizon 16, observation horizon 2, min-max normalization to [-1, 1]. Inference: DDIM sampling with 100 steps to iteratively denoise the action chunk from noise to a clean, executable trajectory. Compared against an MSE baseline: MSE collapses multimodal action distributions to their mean, while the diffusion policy preserves the modes and produces sharper, more realistic actions.

Tech Stack

PythonPyTorchDiffusion ModelsU-Net 1DDDIMImitation Learning