@competesai @nrnagents | Redefining Robustness in Al/Robotics | Insights on Embodied Al, ML Failures & Real-World Deployment | DM for CollabsJoined March 2022
LeRobot v0.6.0 is officially here: Imagine, Evaluate, Improve! 🤖🚀
We are closing the robot learning loop with massive upgrades for the open-source robotics community. From policies that imagine the future to a much leaner installation, here is what is new:
- 🌍 World Models: VLA-JEPA, LingBot-VA, and FastWAM help your policies anticipate the future.
- 👀 VLA Expansion: Welcome GR00T 1.7, MolmoAct2, EO-1, Multitask DiT, and EVO1.
- 🏅 Reward Models API: Track success seamlessly with Robometer and TOPReward.
- 🎯 Unified Evaluation: 6 new simulation benchmarks, all accessible via the lerobot-eval CLI.
- ☁️ And more: lerobot-rollout CLI for DAgger corrections, HF Jobs cloud training, up to 2x faster data loading, GUI - LeLab, many docs improvements
Ready to build the future of robotics? Dive into the full release notes here: huggingface.co/blog/lerobot-r…@ClementDelangue@Thom_Wolf
TL;DR Photometric robustness in VLA models is achievable. One model already proved it.
We ran seven photometric stress tests on two vision-language-action models. Same benchmark, same perturbations, same severity levels. Pi 0.5 held flat. SmolVLA lost ground on nearly every one.
Jack Dorsey just published something that should be required reading for every founder.
The premise: the org chart needs to be replaced entirely. And the argument starts 2,000 years ago.
For thousands of years, every organization on earth has run on the same logic the Roman
Lighting changes constantly: time of day, weather, different rooms, sensor drift. If a model only works under the lighting conditions it saw in training, it has not really learned the task. It has learned one appearance regime.
We put this to the test. Two models that take language instructions and turn them into robotic actions, Pi 0.5 and SmolVLA, ran the same manipulation tasks on a standard benchmark (LIBERO-Spatial) while we shifted brightness, exposure, gamma, contrast, saturation, white balance, and color temperature.
Same geometry, same objects, same tasks. Only appearance changed.
Pi 0.5 barely moved. Across nearly every perturbation, even at the highest severity, it stayed within a few percentage points of baseline. The only measurable dip was contrast, to around 94% of baseline. Not a collapse. A graceful decline.
SmolVLA degraded under nearly every one. Saturation cut performance roughly in half. Brightness produced steady losses. Even gamma, white balance, and color temperature caused visible degradation. And then there was low contrast. SmolVLA went from baseline to near-zero. Not a degradation curve. A complete collapse.
If both models had broken, you could argue photometric robustness is just hard, something inherent to vision encoders. Pi 0.5’s near-total immunity rules that out. Photometric robustness is achievable. SmolVLA’s failure is diagnostic.
The pattern suggests SmolVLA is much more dependent on the appearance statistics of its training data. Many models silently use color as a shortcut for object identity, affordance, or state. When color shifts, those shortcuts break.
By contrast, Pi 0.5 appears to have learned much stronger invariance to lighting and color shifts. Training augmentation is likely part of that story.
The two models do share one vulnerability: low contrast.
Pi 0.5 dips gently. SmolVLA collapses.
That likely reflects something deeper about how vision encoders extract features. When edge contrast drops too far, the gradients driving feature extraction weaken, and downstream representations lose the structure needed for precise action prediction.
Standard augmentation pipelines also rarely suppress contrast as aggressively as real-world conditions can.
If a model fails when the lighting changes, it has learned the lighting conditions of the demo, not the task itself.
Full analysis with interactive visualizations:
x.com/sihing_guppy/s…
Mid/back office bros in finance, your 2026 career pivot chance just got fatter. Nothing says ‘please review our AI tool stack’ like accidentally shipping your entire Claude Code source map to npm 😂
@realmc@CompeteSai@sihing_guppy Got it—bookmarked!
Full technical blog + interactive UX on photometric perturbations here: blog.competesai.com/blog/visual-se…
Pi 0.5's robustness vs SmolVLA's sensitivity is a great reminder: real-world robotics needs vision that separates structure from lighting, not just memorizes
90K Followers 9K FollowingBuilding @liveframe | Distribution + attribution for the attention economy.
I stream here and on Kick: https://t.co/KRbWDCSAE5
162K Followers 6 FollowingStitch by @GoogleLabs turns your ideas into beautiful interface designs, powered by some of the latest Gemini models. Try free of charge.
15K Followers 4K FollowingRobotics researcher & builder working on physical AI, robot identity, and cultural/economic infra for autonomous robots. Prev NASA, Goldman Sachs, Mysten Labs
1K Followers 600 FollowingFight for 1-bit AI🚀, Scalable and Efficient Foundation Models & Deep learning, PhD student, Prev: Research Intern@Microsoft Research
570K Followers 2K FollowingPolyagentmorous ClawFather. Came back from retirement to mess with AI and help a lobster take over the world.
@OpenClaw🦞 + @OpenAI