it's really good that openAI and huggingface were so transparent and collaborative on this safety issue, i think these types of problems are going to be increasingly challenging for any one organization to solve
this is pretty nuts. i think this is the tip of the iceberg for needing to focus on agentic AI safety. models will only get more capable from here, and RL incentivizes getting the right answer at any cost
this is a large reason why we're obsessed with 'intentional design', in other words, making sure models are trained to do the right things for the right reasons
we had a significant security incident during evaluation of our models. we are sharing what we have learned so far. thanks to @huggingface for the partnership on this.
openai.com/index/hugging-…
we're hiring a lead for model training & training research. if you want to make models great at interpretability, and push the frontiers of intentional design, apply here: job-boards.greenhouse.io/goodfire/jobs/…
Today, we are introducing Inkling.
Inkling reasons efficiently across text, image, and audio modalities. We are making the full weights available.
thinkingmachines.ai/news/introduci…
Available today for fine-tuning on Tinker. Play with it in the Inkling Playground. 🧵
> replicate J-space on GLM 5.2
> train a reward model and run RL to reduce hallucinations
> show me how this model makes cancer predictions
Using our platform Silico is like having a team of AI researchers ready to run experiments like these.
Private beta is open now. 🧵 (1/6)
i am consistently amazed by the quality of the experiments that silico can run. excited to kick this off in private beta!
we're also doing grants to select AI safety and academic researchers. DM me if you're working on something interesting
> replicate J-space on GLM 5.2
> train a reward model and run RL to reduce hallucinations
> show me how this model makes cancer predictions
Using our platform Silico is like having a team of AI researchers ready to run experiments like these.
Private beta is open now. 🧵 (1/6)
just welcomed chris earls, cornell engineering professor, to the team!
i'm particularly excited about his work to unlock scientific creativity in frontier AI with interpretability
we've hired several professors at @GoodfireAI because we're investing heavily in foundational research to discover the science of neural networks. if this is work that you're interested in, join us!
I know it might kill everyone, but I can't help loving AI. All my life I wanted to talk to an alien intelligence, but I never thought I'd really get the chance to.
Can LLMs predict the next World Cup champion?
Goodfire partnered with @EternisAI to improve how LLM forecasters use available evidence and manage uncertainty.
We found models were overconfident in their predictions – but probes significantly improved calibration. (1/6)
I heard the same thing from magazine designers when pagemaker came out
Then again when 99designs and logo contests emerged
Not sure if this time is different 🤷
Every founder I meet worries about missing this moment.
The best worry about wasting it. There’s a difference.
It’s a waste to ignore how much the world has changed. It’s a waste to think you can capture value with pure software the same as 5 years ago.
It’s a waste to build
I have an ultrarare, pathogenic variant which is likely to kill me eventually unless biotech advances before then. Standard pathogenicity tools (e.g., REVEL, PolyPhen-2, ClinPred) do not work on this type of variant. What does? @goodfire's EVEE. It uses Evo2, so it's capable of generalization to this variant type, enabling downstream analyses to get disease risk estimates on biobank data
Today, researchers made an important breakthrough in interpretability.
They found "manifolds" in the neural net weights for any concepts in image gen models (SDXL), like the pretzel manifold, and could steer them to generate various kinds of pretzels from the weights directly.
This is well beyond a neat theoretical understanding on how AI models see but gives us a low-compute volume dial to edit the results of a generation.
If models think in shapes, our tools should too.
Our latest research: Block-Sparse Featurizers (BSFs), a new way to find concepts in model activations - using multidimensional “blocks” instead of single directions. (1/9)
Announcing our $130M Series A to build the Open Superintelligence Stack
Led by Radical Ventures, with NVIDIA, Intel Capital, Dell Capital, and existing investors
Train, deploy, and continuously improve your own models using our stack.
Own your intelligence.
We're excited to announce that Resolution has a $160M grant from Coefficient Giving: $108M unconditional, with a further $52M conditional on hiring and compute needs. We'll use it to grow teams across our research portfolio and invest heavily in research automation. 🧵
227 Followers 372 FollowingNo. 10 Downing Street Innovation Fellow | Research Scientist at AISI | Visiting Lecturer at Imperial College London
Working on AI Evaluation and AI for Medicine
0 Followers 135 FollowingResearching AI recommendation altering.
Real website experiments.
Real visibility insights.
Documenting the results in public.
662 Followers 663 FollowingSoftware engineer building AI-powered products solo | Sharing what actually works | Now live: Mind Dojo — train focus, memory & logic | https://t.co/ejJJmGKgRd
2K Followers 3K FollowingAI/ML for proteins, small molecules, and everything in between. 🌈 Past: VP ML @Numerionlabsai (fka Atomwise), UCSF, Columbia, DE Shaw Research, MIT.
25 Followers 203 Following21 // Research Fellow at @law_ai_ // CTO @oxfordunion // Technical AI Security, Alignment and Mech interp researcher // Law Undergraduate @uniofoxford
44 Followers 262 FollowingPronunciation: qwäh | Best trading advice I've gotten: "Any buys today, Dano?" "Coffee." ☕ The wisdom is in knowing when to sit out.
314K Followers 1K FollowingBuilding new things @thinkymachines. Also dabble in robotics at NYU. Cofounded @PyTorch. AI is delicious when it is accessible and open-source.
3K Followers 2K FollowingICA is all we need | AI, Theory, MechInterp, Neuro | assistant professor @CSHL | prev @Stanford @Meta @Google @MPI_IS @ENS_ULM @UCL
153K Followers 1K FollowingSemiAnalysis
Boutique AI Infrastructure Research and Consulting
DMs are open for consulting, quotes, or to talk shop,
Opinions my own
13K Followers 354 FollowingCofounder and Chief Scientist at Resolution. Alignment will be solved, but not necessarily in time. Previously AISI, DeepMind, OpenAI, Google Brain, etc.
56K Followers 384 FollowingProfessor @ Tsinghua, Founder of https://t.co/3IaQ4CI5W3.
AGI, LLM.
“The value of a man should be seen in what he gives and not in what he is able to receive.”―Einstein
95K Followers 924 FollowingOpen model research @ something new.
Prev. co-led Olmo at Ai2.
Contact via email.
Writes @interconnectsai
Wrote The RLHF Book,
🏔️🏃♂️
7K Followers 209 FollowingCo-Founder & CEO @instadeepai Entrepreneur passionate about developing the ML community. @Google For Startups Mentor, @DeepIndaba Steering Committee Member
6K Followers 1K FollowingBehavioral scientist interested in learning and decision-making.
Co-author of Decision Making: a Very Short Introduction https://t.co/uXQze1CBzH
3K Followers 3K FollowingStaff writer @Forbes covering AI and startups. I write The Prompt, a newsletter on AI/ send tips: [email protected] or signal: rashis.17