Instead of training a world model to condition on robot actions, PointZero conditions on partial 3D point tracks. This allows it to learn rich 3D dynamics priors from actionless data!
We found that pretraining on this objective effectively transfers to a variety of robotics applications, outperforming strong application-specific baselines on imitation learning and robot-action-conditioned dynamics modeling.
We trained PointZero only on synthetic data, but the zero-shot real-world dynamics predictions were surprisingly impressive!
Super cool work led by @bardienus
Introducing PointZero—a 3D world model pre-trained without robots.
Dexterous manipulation requires understanding diverse 3D dynamics, but current approaches rely on expensive robot data. So how can we pre-train a 3D dynamics model without it?
PointZero introduces a simple idea:
Introducing PointZero—a 3D world model pre-trained without robots.
Dexterous manipulation requires understanding diverse 3D dynamics, but current approaches rely on expensive robot data. So how can we pre-train a 3D dynamics model without it?
PointZero introduces a simple idea: learning to complete 3D point tracks yields transferable 3D dynamics!
Our model achieves SOTA results on:
🔥 Zero-shot 3D dynamics
🔥 Action-conditioned 3D dynamics (post-training)
🔥 Imitation learning (post-training)
Details and links 👇 1/6
Humans learn by interacting with the world: How can we teach robots to similarly adapt to ambiguous dynamics? Introducing Wiggle and Go! #CoRL2026. We find that a specialized model can learn rope properties from just one interaction, and complete complex dynamic tasks.
The robot wiggles a rope once, figures out how it behaves, then goes, using the estimated dynamics to execute the task. No real-world training data. No retries. #robotics#robotlearning
We also tested ModAR on real-world bimanual manipulation tasks:
ModAR beats baselines and also shows that it can improve its performance by training on human videos,
using videos from both the same setting (ID) and a public dataset (OOD).
🧵 [6/8]
We also fine-tune Flex-π, a recent multimodal WAM that initializes from a pretrained video model, on the same data. Despite ~200× fewer parameters (30.1M vs 6B) and ~20× fewer training FLOPs, ModAR achieves a slightly higher average success rate: 75% vs 72%.
We think this demonstrates a promising new direction:
scaling from-scratch WAMs with structured representations that transfer better to manipulation,
rather than relying on RGB-centric video models.
🧵 [5/8]
Next: how should future prediction and action prediction be coupled?
• Unified: denoise everything together
• Disjoint: denoise futures and actions independently; skip generating futures at test time
• ModAR: denoise modalities sequentially, then actions
ModAR wins at every data scale – and benefits most from actionless data.
🧵 [4/8]
First: what modalities should a WAM imagine in? Does adding more modalities help?
Result: Progressively adding Tracks, DINO, and depth improves performance, but adding RGB on top does not.
🧵 [3/8]
📑 arXiv: arxiv.org/abs/2609.17524
🌐 Project Page: adamhung60.github.io/ModAR/
Why generate modalities one at a time?
The idea is simple: predict the future in easier, more structured forms first, then use those predictions to help generate harder, more detailed modalities.
To test this, we conduct a controlled study across six RoboTwin tasks, varying the predicted modalities, WAM formulation, and amount of actionless training data.
We ask two main questions: which modalities to predict, and how to couple future prediction with action prediction.
🧵 [2/8]
World-action models typically imagine the future in RGB – but are pixels really the right representation for robotics?
Our bet is no: RGB spends capacity on fine-grained details and variation that are often irrelevant for robot policies.
DINO features, point tracks, and depth capture more useful features like semantics, motion, and geometry.
But no single modality captures everything — how can we effectively combine them?
We introduce ✨ModAR✨, which predicts the future one modality at a time, with each prediction informing the next. We train from scratch and find that this formulation performs best.
Our 30.1M scratch-trained model even outperforms a 6B video-model-initialized model finetuned on the same data!
🧵 [1/8]
Today we are introducing Atlas! An autoregressive and multimodal DiT for generating image and depth frames. Atlas is a state-of-the-art 3D reconstruction model and outperforms all open-source models. @HaoZhang623 and I had lots of fun pushing the 3D capability of this model; can't wait to see you build with it!
Introducing Atlas:
The world's first multimodal world model that generates image and video frames with pixel-perfect camera control and reconstructs them in 3D.
Model the world, move the camera, and simulate space & time.
[1/6] Can one model fix rendering artifacts from any 3D representation? And what does it even mean to "fix" a rendering?
At #ECCV2026, we present FixAnything (fix-anything.github.io): a single generalist video model that refines 3DGS/mesh/sparse point-cloud renderings into photorealistic, 3D-consistent videos.
Can robot learn manipulation skill with one object and perform it on a completely different one without any new demonstrations?
Excited to introduce SemAnCorr, a framework for dense correspondence and manipulation skill transfer across diverse objects.
semancorr.github.io
How can we get robot hands to “hear” slip and contact through microphones, and react to them?
We’re excited to share VibeAct, an approach that uses piezoelectric microphones embedded in robot fingertips to estimate contact and slip, then learns reactive policies from this tactile feedback!
vibeact.github.io
Excited to share the first paper of my PhD!
If you’ve ever tried to control a VLA via natural language, you know it rarely does what it is told. 🗣️ We introduce a multi-stage pipeline for training a Language Feedback Policy (LFP) to steer a VLA in-the-loop.
Introducing Modality Forcing, a recipe for post-training T2I models for SOTA RGB-Depth generation!
Text-to-image (T2I) models learn rich representations of the spatial world.
How do we build on this prior for high-quality depth generation?
modality-forcing.github.io
🧵 [1/6]
264 Followers 2K FollowingLLM Post-Training/ Recursive Self-Improvement for Reasoning
Co-founder/ MTS @ReasonCoreAI
PhD @UCLA, Bachelors @iitbombay
Led a team of ~500 PhD researchers
627 Followers 7K FollowingMachine learning engineer working on signal processing, NLP and (some) computer vision.
Trying to make the most of my measley 3090 and 3 Sparks.
86 Followers 152 FollowingLead vision engineer at @1x_tech, former co-founder CTO at https://t.co/qx6aarbrPs interested in learning based robotics, embedded systems, vision etc.
12K Followers 10K FollowingProfessor @UAM_Madrid
Writing on what matters in AI for Science: https://t.co/sNvjOIcqgU
"IA para estudiantes de ciencias": https://t.co/Zs5o2gszjj
86 Followers 152 FollowingLead vision engineer at @1x_tech, former co-founder CTO at https://t.co/qx6aarbrPs interested in learning based robotics, embedded systems, vision etc.
3K Followers 355 Followingai & robotics @Amazon | writing about research & life.
prev: @meta @nvidia @UCBerkeley @MIT
🌈 she/her/hers
I like capybaras :D
665 Followers 863 FollowingWorking on robot learning and foundation models @nvidia. PhD from @CMU_Robotics. He/Him. Views are my own and don't represent my employer.
105K Followers 742 FollowingVP Digital Human Research, Epic Games. Emeritus Director, Max Planck Institute for Intelligent Systems (@MPI_IS). Opinions are my own.
2K Followers 1K FollowingRobotics researcher at @NVIDIA. Ph.D. @BrownCSDept 🐻. Prev: @ColumbiaCompSci, @rai_inst, @Amazon, @Honda. Opinions are my own and do not represent NVIDIA.
8K Followers 2K FollowingAssistant Professor @CMU_Robotics @SCSatCMU, leading the @LeCARLab. @Amazon Scholar @ FAR. I build generalist robots with agility. Ph.D. from @Caltech.