GPWM: Generative Particle World Models
for 3D Prediction and Planning
Anonymous authors · paper under double-blind review
Abstract
We introduce GPWM, a generative world model that learns 3D physical dynamics directly from particle trajectories. Its shared representation and architecture covers rigid, deformable, cloth, and granular materials, in both passive physics and robotic interaction. GPWM is an object-aware multi-resolution transformer that combines flow matching with diffusion forcing to enable streaming generation far beyond the training horizon. We train and evaluate GPWM on ParticleWorlds, a 25.5-million-frame corpus spanning diverse geometries, physical properties, interactions, and robot embodiments.
When trained within individual domains, GPWM reduces prediction error by 29–53% relative to the strongest applicable prior learned-dynamics models. Going beyond the domain-specific setting of prior methods, a single Unified GPWM jointly models heterogeneous physical regimes within one set of weights. GPWM further generalizes to unseen geometries, physical properties, object counts, and rollout horizons.
Beyond prediction, GPWM enables contact-rich planning. An object-space goal reward is backpropagated through the learned interaction dynamics to the latent noise, steering actuator motion and the contacts it establishes. Executed open-loop in simulation, the resulting plans outperform unguided generation and planning with prior learned dynamics models on pushing, picking, and granular manipulation.
Method
An object-aware, multi-resolution spatio-temporal transformer
Given the observed frames, GPWM models a distribution over the next frames, optionally conditioned on material properties and prescribed actuator motion.
The network progressively compresses dense particles into a small set of anchor tokens, performs most global interaction modeling at coarse resolution, and propagates the resulting features back to the original particles.
Dataset
ParticleWorlds: 25.5 million frames in six domains
Simulated trajectories spanning rigid objects, elastic and elastoplastic solids, granular materials, and cloth, covering passive dynamics and robotic manipulation with rods, parallel-jaw grippers, and articulated hands.
Streaming
Streaming long-horizon generation
Near-future frames are progressively resolved, while farther-future frames remain at higher noise levels and continue to be revised. Once the earliest frames are fully denoised, they are committed, the prediction window advances, and fresh Gaussian noise is appended at the far end of the horizon.
Comparison
Comparison with learned 3D dynamics models
GNS, a graph-based particle simulator; PGND, a hybrid particle–grid model; PointWorld, a point-cloud world model; and PhysCtrl, a generative model of physical dynamics. All methods use the same evaluation trajectories and available actuator conditioning.
Planning
Reward-guided planning
A differentiable reward specifies the desired final configuration of the passive objects. With the world model fixed, the object-space reward guides actuator motion through the learned joint dynamics, without requiring a reference actuator trajectory or training a goal-specific policy.
Real world
Predictions from a single image
Previously unseen object geometries reconstructed from single RGB images, with their dynamics predicted by the Free Fall specialist. The resulting trajectories are rendered back into the scene.
Try it
Drop a ball anywhere on the desk
Click the table to drop an elastic or a sand ball.
Results
Prediction and planning, in numbers
BibTeX
@misc{anonymous2026gpwm,
title = {{GPWM}: Generative Particle World Models for {3D} Prediction and Planning},
author = {Anonymous},
year = {2026},
note = {Under review}
}