A robot can follow a reasonable interpretation of an instruction and still fail the benchmark.
Meet BenchMend: we use GPT-6-Astra to uncover instruction–verifier gaps in robotics benchmarks, then repair the wording while preserving success checkers and recorded demonstrations.
Excited to share @AlanZhang0928's BenchMend!
Robots may fail a benchmark not because of wrong actions, but a misaligned instruction-verifier. GPT6 Astra can repair this with minimal change, pointing toward benchmarks that uncover their failure modes and iteratively self-improve.
A robot can follow a reasonable interpretation of an instruction and still fail the benchmark.
Meet BenchMend: we use GPT-6-Astra to uncover instruction–verifier gaps in robotics benchmarks, then repair the wording while preserving success checkers and recorded demonstrations.
Today we are introducing Dyna-2, a world-action model pre-trained on one million hours of human video. At this scale, for the first time, we discovered several new scaling laws:
• world-action models exhibit scaling law on human data across four orders of magnitude, from 1000 to 1,000,000 hours,
• this human data scaling law implied a scaling law on never seen robot data,
• both data and objective matter; world modeling and scaling on video data are essential for cross-embodiment scaling transfer to emerge
🧵
Thanks Shuran! Really appreciate your support. Website: handroid.org Also excited to see the amazing and inspiring Transformer Transformer for co-design transformer-transformer.github.io
Thanks Haochen! It was a pleasure collaborating with you as well. Your insights were incredibly valuable and look forward to pushing this direction even further together.
This is exactly the kind of "transformer" I imagine. Creatively extending the otherwise limited workspace of a robot arm in a way that just makes sense. A perfect example of a brilliant idea + amazing execution. Enjoyed collaborating with the team a lot!
🤖 A robot that transforms between humanoid and dexterous hand?
Introducing Handroid, a reconfigurable robot with a shared 27-DoF body 🖐️→🚶→🖐️:
• The same joints, different roles
• Expanded robot task space
• Shared sensing and control interfaces
handroid.org
Humanoids shouldn’t have to stop before they act.
CoorDex enables humanoids to walk and manipulate simultaneously by coordinating body and dexhand motion priors in latent space.
AnyBody enables coordinated whole-body control from sparse keypoints, while adapting to environment.
Sharing my CVPR 2026 talk from the Vision for Intelligent Task Assistants workshop: "From Perception to Agency: The Cognitive Stack for Video Task Assistants."
It covers our SVI-Bench project (svi-bench.github.io) plus a video+robotics project we'll release soon.
w/
What if humanoids could manipulate without stopping, using a high-DoF dexterous hand?
🚨We introduce CoorDex, a learning pipeline for continuous dexterous humanoid loco-manipulation — coordinating whole-body motion and high-DoF finger control while the robot is still moving.
This enables a Unitree G1 with a 20-DoF WUJI dexterous hand to:
🤖 grasp and carry a bottle while walking
🚪 open a fridge while stepping backward
🔄 pick up a cube and turn 180°
🔗Project page: skevinci.github.io/coordex/
📜Paper: arxiv.org/abs/2606.23680
🧵(1/n)
🔥 AutoResearchClaw tech report + v0.5.0 just dropped.
12,300+⭐ on GitHub. Two big additions this release:
🧪 1/ Domain-Expert Agents in the experiment stage: Specialized agents for high-energy physics, biology, and more. Real domain tools + knowledge plugged in — not a generic LLM pretending to run experiments.
📊 2/ ARC-Bench
A 55-topic benchmark across ML, HEP, quantum physics, biology, and statistics. One of the broadest cross-disciplinary evaluations for autonomous research ever released.
🏆 The numbers:
→ Beats AI Scientist v2 by 54.7% on ARC-Bench
→ 7-mode HITL (human-in-the-loop) ablation: targeted intervention > full autonomy OR exhaustive oversight.
The thesis (still): real research isn't a pipeline. Hypotheses fail. Lessons compound. AutoResearchClaw is a research amplifier — not a paper generator.
📄 Tech report: arxiv.org/abs/2605.20025
💻 Code: github.com/aiming-lab/Aut…
Thanks @itsJiaqiLiu and @StephenQS0710 who lead the work and all other contributors @HaonianJi, @lillianwei423, @XinyeYee, @richardxp888, @HaoqinT, @Xinyu2ML, @WeitongZhang, @jiahengzhang96, @LINJIEFUN, @linjunz_stat, @yuyinzhou_cs, @CaimingXiong, @james_y_zou, @ZhengBerkeley, @cihangxie, @dingmyu
Everyone's excited about Karpathy's autoresearch that automates the experiment loop.
We automated the whole damn thing. 🦞
Meet AutoResearchClaw: one message in, full conference paper out. Real experiments. Real citations. Real code. No human in the loop.
One message in →
🤖Can robots play Texas Hold’em? 🤔
We introduce DexHoldem, a real-world benchmark for dexterous manipulation and embodied agents.
After fine-tuning leading VLAs and pairing them with frontier agents, we found:
Robots can move the chips — sometimes.
Agents can read the table state — sometimes.
Unfortunately, poker requires both to happen correctly in the same hand.
So no, robots are not ready to take your money at the casino — yet.
More below 👇
🚀Introducing DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLAs. By formulating latent visual foresight as a differentiable structural bottleneck, DIAL strictly grounds motor control in the VLM’s intent. This architecture alleviates representation collapse during joint optimization, enabling the model to absorb cross-embodiment human data for robust generalization, while mastering complex multi-stage coordination in the real world.
🚀Paper, code, and project page released! Fully open-sourced!
Paper: arxiv.org/pdf/2603.29844
Project page: xpeng-robotics.github.io/dial/
Code: github.com/xpeng-robotics…
Congrats to @ChenYi041699 and many thanks to our awesome collaborators @tttoaster_@dingmyu@ge_yixiao
🚀 AutoResearchClaw v0.4.0 is here — almost 10K⭐ in just over 2 weeks!
Now supporting both fully autonomous AND human-AI co-pilot modes — you choose your level of involvement.
What's new:
🤝 6 intervention modes — full-auto, gate-only, checkpoint, step-by-step, co-pilot, and custom. Same powerful 23-stage pipeline, your level of control.
🧪 Idea Workshop — brainstorm and refine hypotheses with AI before committing to a direction
📊 Baseline Navigator — review and customize experiment designs before execution
✍️ Paper Co-Writer — draft papers section-by-section, collaboratively
🧠 SmartPause — the system learns when to pause and ask for your input based on confidence levels
💰 Cost Guardrails — budget alerts at 50/80/100% so you never get surprised
🔀 Pipeline Branching — explore multiple hypotheses in parallel, compare, and merge the best
Want full automation? It still does that. Want to stay in the loop? Now you can, at exactly the granularity you want.
Try it 👉: github.com/aiming-lab/Aut…
Kudos to the team @JiaqiLiu835914, @richardxp888, @lillianwei423, @StephenQS0710, @Xinyu2ML, @HaoqinT, @jiahengzhang96, @yuyinzhou_cs, @ZhengBerkeley, @cihangxie, @dingmyu, etc.
Do Vision-Language-Action Models truly follow your language instructions?
We present When Vision Overrides Language: Evaluating and Mitigating Counterfactual Failures in VLAs. They promise to ground language instructions in robot control, yet in practice, often fail to follow language faithfully.
📄 Paper: arxiv.org/abs/2602.17659
🌐 Project: vla-va.github.io
💡 Highlights
Vision shortcuts and counterfactual failures. When given instructions that lack strong scene-specific supervision, they default to well-learned scene-specific behaviors regardless of language intent.
Counterfactual benchmark. We introduce LIBERO-CF, the first counterfactual benchmark for evaluating language following in VLAs. Our evaluation reveals that counterfactual failures are prevalent yet underexplored across state-of-the-art VLAs.
Our solution. We propose Counterfactual Action Guidance (CAG), a simple plug-and-play dual-branch inference scheme that strengthens language conditioning without changing pretrained VLA architectures or weights.
Experiments. CAG is effective across multiple dimensions of language grounding, consistently improving both language grounding and task success on under-observed tasks.
#VLA#Robotics#Vision#Language
Can we build a universal brain for all dexterous robot hands?
Zhenyu Wei, Yunchao Yao, and Mingyu Ding from University of North Carolina at Chapel Hill just tackled this!
By creating a "canonical representation," they translate all kinds of dexterous robot hands into a single, unified description and control language.
This allows a single AI policy to understand and control them all. The result: policies that instantly generalize to any new robot hand design, achieving an 81.9% zero-shot success rate on unseen hands and opening the door to universal dexterous manipulation.
One Hand to Rule Them All: Canonical Representations for Unified Dexterous Manipulation
Project: zhenyuwei2003.github.io/OHRA/
Paper: arxiv.org/abs/2602.16712
Code: github.com/zhenyuwei2003/…
Our report: mp.weixin.qq.com/s/cp15BVTkxkZM…
📬 #PapersAccepted by Jiqizhixin
Found an impersonation account @mingyding pretending to be me.
I only have one account. Please do not interact with the fake account and help report it if possible, thanks!
425 Followers 2K FollowingCEO & Founder @origin_autonomy | Building Time Machines for Construction!
Ex-Founder - Perpule (Acq by Amazon) | AI x Robotics | ACM-ICPC World Finalist
21 Followers 144 FollowingI am a first-year PhD student at SFU GRUVI lab, working with Professor Manolis Savva. I'm broadly interested in computer graphics and robotics.
2 Followers 227 FollowingPh.d in @sjtu1896, intern in BigAI, @SharpaRobotics
Supervised by @siyuanhuang95, @ChuanWen15
Research focus on dexterous manipulation & computer vision
283 Followers 209 FollowingHi! I am a final-year undergraduate student at @sjtu1896. I am interested in Robot Learning and Humanoid Robots. Website: https://t.co/S5lYcsqWhP
5K Followers 252 FollowingFounder @AetherLab_AI
Assistant Professor @HDSIUCSD @UCSanDiego
Causal World Model, Causality-driven Agentic System for the next AI paradigm
268 Followers 155 FollowingPh.D. student. at the University of Hong Kong, advised by prof. Yi Ma
Graduated from IIIS Tsinghua University 2024 (Yao Class)
11K Followers 579 FollowingCo-Founder of SceniX (now part of @theworldlabs) | Assistant Professor @Columbia @ColumbiaCompSci | Former postdoc @Stanford @StanfordSVL | PhD @MIT_CSAIL
3K Followers 1K FollowingCS Ph.D. student @Columbia & Research Scientist @NVIDIARobotic | Prev. Meta FAIR Embodied AI, Boston Dynamics AI Institute, Google X #Vision #Robotics #Learning
6K Followers 70 FollowingProf @Stanford, Distinguished Research Scientist and AV research lead @nvidia. PhD from @MITAeroAstro. Robotics, autonomous systems, AI. Opinions are my own.
5K Followers 376 FollowingI am an associate professor at UT Austin. I do research at the intersection of computer graphics, computer vision, and machine learning.
2K Followers 981 FollowingAssistant Professor @HKUST | AI for Creativity | ex @Stanford @Meta @RealityLabs @UofT | Works #ControlNet #AnimateDiff @cveu_workshop
22K Followers 1K FollowingProfessor @ucsantabarbara. Head of Research @SimularAI. Director @ucsbcrml @UCSB_AI. Build the Science of AI Agents. AI for Humanity in the long run.
129K Followers 23K FollowingWelcome to the U.S. Postal Service official Social Customer Response site. Hours of Operation: M-F: 8am-10pm ET. Sat: 9:30am-6:00pm ET. Closed Sunday.
359K Followers 121 FollowingThe official X account of the United States Postal Service, managed by the Social Media staff at USPS HQ. For customer service, please follow and DM @USPSHelp.
528K Followers 6K FollowingCo-founder & CEO of Discovery Loop. Former Chief Scientist, Google. Helped build many Google products, TPUs, Gemini, TensorFlow, MapReduce, Bigtable, ...
2K Followers 2K FollowingAssociate Professor in CS at UNC-Chapel Hill. Research crosses visualization, cognitive science, data science, and MR. Mom, hockey player, golfer, & crossfitter