Our takeaway: multimodal models possess strong reasoning abilities, and VISTA can unlock their potential to solve tasks across diverse interactive environments. Check out the paper for more details!
Our 🔭VISTA🔭 paper is now on arXiv, with more ablations and results!
We introduced VISTA in our Aug 5 blog post. The paper further evaluates it on visual games and reasoning tasks beyond ARC-AGI-3 🧵
Paper: arxiv.org/abs/2610.02200
Code: github.com/joshhhhhan/VIS…
Previous blog post: vista-research.github.io
How does a model understand and explore a world it has never seen?
We introduce 🔭VISTA🔭, a visual harness that gives a VLM long-horizon vision for reasoning in an interactive world. With Claude Opus 5.0, it reaches 100% RHAE on @arcprize's ARC-AGI-3, perfectly solving all 25
🚀One of the biggest questions for AI agents is whether they can continue expanding their capabilities after deployment.
Deployment brings the experience needed to keep improving. As @ilyasut has argued, future intelligent systems should learn from deployment.
But this requires more than a new learning algorithm. It requires turning the serving stack itself into a learning layer: collecting live experience, turning it into updates, and safely bringing those updates back into serving.
That’s why we built Reef.
Reef is open-source infrastructure for continuously evolving agents at live deployment. To our knowledge, it is the first open-source infrastructure designed to evolve both model weights and the agent harness from deployment experience.
Not just weights, but also prompts, memory, skills, tools, and orchestration.
Reef already supports:
🧠Model evolution: SAO, TTT-Discover, OpenClaw-RL, with more recipes coming.
🛠️Harness evolution: SkillClaw, Meta-Harness, and a general harness-evolution engine built on Cordis, with native support for pi @pidotdev , OpenCode @opencode , and more harnesses coming.
With Reef, inference is no longer the end of the pipeline. It becomes part of a continual loop:
serve → learn → evolve → serve again.
Reef is fully open source. We would love you to try it, build on it, and tell us what is missing!
⭐ GitHub: github.com/Human-Agent-So….
💬 Discord: discord.com/invite/5y8e5f9….
#AgenticAI#llms
VISTA is also not limited to 2D games. The same approach works when the same world is rendered as a 3D scene, or serialized into a 1D text grid, opening up more opportunities to generalize VISTA to other complex interactive tasks!
How does a model understand and explore a world it has never seen?
We introduce 🔭VISTA🔭, a visual harness that gives a VLM long-horizon vision for reasoning in an interactive world. With Claude Opus 5.0, it reaches 100% RHAE on @arcprize's ARC-AGI-3, perfectly solving all 25 public games.
Blog post: vista-research.github.io 🧵
Language is discrete. Language models don’t have to be.
🧚Introducing ELF🧚♀️: Embedded Language Flows—a class of diffusion models in continuous embedding space based on continuous-time Flow Matching 🧵
25 Followers 448 FollowingFounder @ Veyra AI
Building LLM answer-eval infrastructure for brand inclusion & ranking
Ex-RecSys / large-scale ML systems
Obsessed with decision stability & m
263 Followers 3K Followingcomputer vision, crossfit-er, salsero, previously augmented reality at Microsoft and autonomous vehicles at NVIDIA. Now, future communication at Google Beam.
37 Followers 207 FollowingCS PhD @USC | student researcher @GoogleDeepmind. Teaching agents to build the 3D world via writing procedural code. Views are my own.
8K Followers 1K FollowingAssistant Professor at @NUScomputing | Robotics & AI PhD @uwcse | Host of @RoboPapers | Ex-@allen_ai robotics lead, @NVIDIA | Magician |
my opinion is my alone.
2K Followers 788 Followingcs phd @stanford | interned @nvidia, seed | prev. @eth | I work on long context, continual learning, generative models, and infra | opinions are my own
5K Followers 1K FollowingPostdoc @berkeley_ai, building Impossible. Prev. @MPI_IS @AdobeResearch @uni_tue Interested in recreating the physical world and the intelligence to do so.
528K Followers 6K FollowingCo-founder & CEO of Discovery Loop. Former Chief Scientist, Google. Helped build many Google products, TPUs, Gemini, TensorFlow, MapReduce, Bigtable, ...
12K Followers 199 FollowingAssistant Prof at Stanford CS, member of @stanfordnlp and statsml groups; Formerly at Microsoft / postdoc at Stanford CS / Stats.