Diffusion Policy for Push-T
Re-implemented Diffusion Policy for the Push-T benchmark with a Conditional U-Net 1D, cosine noise schedule, EMA, and DDIM sampling, preserving multimodal action distributions instead of collapsing to a mean.
Details
Re-implemented Diffusion Policy (Chi et al., RSS 2023) for the Push-T benchmark: instead of predicting actions directly, the model learns to denoise a noisy action horizon conditioned on recent observations, which makes it well suited to multimodal behavior where the same state may admit several valid actions.
Architecture: Conditional U-Net 1D with down dimensions [256, 512, 1024], kernel size 5, and FiLM conditioning. Global conditioning is built from the first n_obs_steps observations.
Training: epsilon prediction with MSE loss, cosine noise schedule via diffusers.DDPMScheduler, EMA (inv_gamma=1.0, power=0.75), and AdamW (lr=1e-4, betas=(0.95, 0.999), weight_decay=1e-6). Horizon 16, observation horizon 2, min-max normalization to [-1, 1].
Inference: DDIM sampling with 100 steps to iteratively denoise the action chunk from noise to a clean, executable trajectory.
Compared against an MSE baseline: MSE collapses multimodal action distributions to their mean, while the diffusion policy preserves the modes and produces sharper, more realistic actions.
Tech Stack
PythonPyTorchDiffusion ModelsU-Net 1DDDIMImitation Learning