Javier Marin @jamarinval
Applied ML: verification & grounding for LLMs. Maintainer of Groundlens and Otwin. RAG and agentic evaluation. Papers on hallucination geometry and Physical AI groundlens.dev Madrid Joined August 2010-
Tweets204
-
Followers388
-
Following563
-
Likes87
"AI safety has a friendly-fire problem. The technology that proves where AI text came from is weakening the technology that checks whether it is true." @jamarinval presents an incisive analysis of the link between LLM watermarks and hallucinations towardsdatascience.com/hallucinations…
Read the full paper. Congratulations for the work. For me, the proposal/consequence separation is the cleanest part of the design. I found that the world acting as the latent space for invention is almost literally Kauffman's adjacent possible theory (TAP), with every installed artifact enlarging what the next agent can build. I'm curious about the control parameter. In your atoms-in-a-box analogy, superconductivity only shows up in a narrow regime. Is there an equivalent here, a disturbance rate or artifact density beyond which the world stops being memory and becomes noise? Your Fig. 15A phase portrait seems to hint at it.
"A threshold tells you an answer is bad, and when it is wrong it costs you. A list of words only tells you where to look, and when it is wrong it costs you a glance. One check takes five minutes and gets skipped. The other takes thirty seconds and gets done." @jamarinval shares an incisive reflection on AI hallucinations and how we try to prevent them. towardsdatascience.com/ten-is-not-a-h…
@SvenUrbanSci @TDataScience Indeed. This was the idea behind the post.
In a thoughtful analysis, @jamarinval reflects on the links between LLM watermarks, hallucinations, and the slippery notion of authenticity in the age of AI. towardsdatascience.com/hallucinations…
"A bank runs a RAG system behind its customer-service chat. A user asks about an invoice. The retrieved document says the total due is 10,000 dollars. The chatbot tells the client that the total is 1,000 dollars." @jamarinval unpacks a RAG pain point that seems simple to solve — until you try to. towardsdatascience.com/ten-is-not-a-h…
Asked Claude.ai for a number that isn't in the doc. It refused. The grounding check flagged it anyway, not because the answer was wrong, but because it didn't come from the source. Provenance, not truth.
How can you ensure that a model grading its own work does it in a robust and cost-effective manner? @jamarinval digs into an emerging conundrum of agentic-AI workflows. towardsdatascience.com/design-loops-n…
The daily "lying" is the case for not trusting the model's own account of itself. If it asserts false things confidently, self-checks inherit the same failure - you need a check anchored to the source, outside the model. Detecting ungrounded answers is tractable; getting the model to self-certify isn't.
The legibility-erosion vs unfaithful cut is the one that matters - but from the outside you often can't tell which regime you're in without ground truth you don't have. Unfaithfulness scores only cover cases where a faithfulness test is constructible. So if safety leans on the CoT, it leans on a proxy you can only audit where you already know the answer - which argues for anchoring the check outside the trace.
Every AI headline is really an energy headline now - and storage is where the grid absorbs the shock. The hard part isn't installing more batteries. It's operating them: predicting failures before they cascade and optimizing dispatch in real time, on assets that degrade in ways a static model never captures. That's exactly where a physics-informed digital twin earns its keep. I'll be on a panel on this at the IEEE PES General Meeting in Montreal - "AI-powered Digital Twins for Grid-Scale Energy Storage" (Wed July 22, 1pm). If you work on storage, grid ops, or AI-driven load, come say hi.
This is the strongest case I've seen for keeping the "don't believe this" outside the model. If a document that explicitly flags a claim as false still gets absorbed as true, the correction can't live in the model's own judgment - grounding has to be checked against the source, not delegated to the model's read of it.
One hard part of model forensics: the "why" often isn't localized. When you compare how a model processes a correct vs. an incorrect answer, the discriminative signal is spread across the trajectory - not sitting in any single layer or component. The "why" may need attribution across the whole path, not one circuit. Wrote it up here: arxiv.org/abs/2603.13259
I just published Do You Know What You Want? medium.com/p/do-you-know-…
It’s clear that current autoregressive models do not equal human intelligence. Everybody knows that. But I do not understand why so much people acts like “well, AI are only good pattern matching systems”. We should remember that our universe started with hydrogen + helium and end up writing poetry.
Can Neural Networks "innovate" during its training? towardsdatascience.com/i-measured-neu…
@chrisalbon @goodfellow_ian I think many of us owe a lot to @goodfellow_ian
Just benchmarked Claude, ChatGPT, Gemini & Grok against each other. Here's what I learned that public leaderboards don't show. The irony of AI benchmarking is this: benchmarks have driven massive progress in the field, but they're almost useless for choosing the right model for your business. We ran the a propietary framework across 4 leading providers, measuring: ✓ Consistency (behavioral reliability across multiple trials) ✓ Performance (quality on real-world business tasks) ✓ Coordination effects (multi-step workflow performance) The results? There's no winner. Or rather, there are four different winners: 🔵 Claude excels at consistency (0.94 score) — if you need auditability, compliance, and predictable behavior, Claude is your model. Regulatory bodies will love it. Cost premium is worth it for decision-critical workflows. 🟡 Grok maximizes performance (88/100 score) — pure output quality. Creative problem-solving, complex analysis, "give me your best answer" tasks. Trades consistency for ceiling height. 🟢 Gemini balances both — neither specialized nor weak. Great if you have diverse workloads and want to minimize switching risk. 🔴 ChatGPT holds the middle — reliable across domains, broad ecosystem, trusted integrations. Your vendor choice matters way less than testing on your actual data. We found that cost-quality tradeoffs, edge case handling, and degradation under load are completely specific to each use case. If you're deployed to production with AI: spend 10-20 hours self-benchmarking. It's ~0.1% of annual AI infrastructure cost and one of the highest-ROI investments you can make. The benchmark that matters most? The one you run yourself. 👉 For more details about the pilot test you can DM me.
AI for Africa 🤖 @Ibitoye198385
2K Followers 5K Following AI • Tech •Crypto. Forex. Useful tools, opportunities & ideas to work smarter + earn online. 🌍 Africa → Global
Marijan Bogdan @PulizInfo
2 Followers 65 Following Diversed,will be known after direction of succession is successfuly applied.
Aban Hasan @AbanHResearch
55 Followers 469 Following Epistemology (Phil) + AI (Applied and Theoretical Alignment) Research Private: @abantheseeker
novinho @022_novinho20
570 Followers 3K Following
KYS @BeastmodeBoothe
1 Followers 15 Following
Frieda @U41La1j7ny19R
158 Followers 6K Following She shines not because she wants to be seen, but because she cannot help it.
Arwen @DachMarian88131
219 Followers 7K Following
Lorenz @Lorenzifix
983 Followers 979 Following
jcamdr @jcamdr70
1K Followers 2K Following The joint system’s can exceed that of any isolated human brain, delivering richer consciousness than evolution prepared us for.
Claudia @jdzK6U86QOraI
106 Followers 3K Following
Damian @Damian827689
6 Followers 180 Following
NoQuienTúQuieres @NoQuienTQuiere1
301 Followers 3K Following Veloz hacia su sino. «¿Por qué se escuenden? ¿Por qué no se mostran?» «En prueba de la verdad de mis palabras os presento un testigo irrecusable, mi pobreza.»
Frank Alcantara @frankalcantara
360 Followers 3K Following Pai, marido, engenheiro, professor e sempre estudante. Father, husband, engineer, teacher, and always a student.
orhiefu @orhiefu75641
7 Followers 152 Following
Olga Kurylenko @qZhybv6K9I2798
429 Followers 4K Following
Jolene @f7sc7931Dyrj463
39 Followers 987 Following The most alluring thing a woman can have is confidence.
StephaniePollitt @Qy2FX0t7sZij5
40 Followers 950 Following
VixenScarlettNelson @e2YcGHvSdK4sy
21 Followers 1K Following Rising above the challenges Making magic happen
Rain Lux @Ra_in_Lux
42 Followers 81 Following A creator of opportunities for ideas, a shaper of the future. A collaborator.
Thalia Rutherford @Tugi1771228
12 Followers 825 Following
Magnivel Internationa... @magnivel
78 Followers 3K Following Magnivel International is an Open Access publisher and international scientific conferences and expo Organizer.
ShirleyMathilda @i281fSRfC45EQD
59 Followers 1K Following
Adesanmi Emmanuel @Sanmyteee
417 Followers 960 Following
Roger Koppl 🗽 @Roger_Koppl
1K Followers 320 Following War bad, peace good. Deportation bad, immigration good. Tariffs bad, free trade good. Tyranny bad, liberty good.
Mentis 🇦🇺 @adam_x_mentis
9K Followers 6K Following Come Vibe Code with me! ✨ Join જ⁀➴@vibeacademy NOW! AI Product Dev | Research Analyst | Tech Wizard | Storyteller | Also cofounder @gaiinsights
Elpsy(Lψ) @haanshin
12 Followers 19 Following In reality I am silent, but here I reveal endless worlds. In the age of AI and physics, I write from the fracture of time.
A man with broken win... @afmothernature
158 Followers 1K Following
Pavel V., MD PhD @SeshatCZ
426 Followers 6K Following Altruistic individualist. Senior Medical Oncologist. IT Advanced User.
Muhammad Usama Saleem @usamasaleem_
39 Followers 284 Following applied scientist @amazon | ex. @Google | ph.d. in multimodal perception & generation
Satora @Satora_ai
653 Followers 5K Following Your go-to product brain. Get started at https://t.co/A0uD304cwF 🪐
Fahad Javed @Fahad_Javed_761
1 Followers 139 Following
Nina @Moroz4240
92 Followers 7K Following I love learning new knowledge. In life, I believe that kindness and smile can bring good changes, and I also hope to meet more interesting people.
Tima @timafanvue
2K Followers 7K Following 20 | Lazy... but I have big milkers Anyway... Very new to onlyf so idk what im doing :) xdd I ONLY REPLY HERE ⬇️
LouisePalmer @lGb9E9tu4wJJJZ
34 Followers 1K Following
SpringGill @8BS1ofe1K3Khu
31 Followers 1K Following
Matthias Schmidt @eurofounder
108K Followers 234 Following Founder based in the EU • Building GDPR-compliant startups • 7 years in, €7k MRR
Sven Nachtzeit @SvenUrbanSci
68 Followers 270 Following „Die einzige Möglichkeit, die Zukunft vorherzusagen, besteht darin, die Macht zu haben, sie zu gestalten.“ – Eric Hoffer. Denker an der Schnittstelle von Wirtsc
Christos E. Athanasio... @christos_edward
4K Followers 3K Following building an eco-future | mechanics & AI & material reuse | Asst Prof @GeorgiaTech
Hyunwoo Yuk @HyunwooYuk
4K Followers 285 Following Scientist's mind, engineer's hand, entrepreneur's heart | Full-time rabbit lover | Founder & CTO of SanaHeal, Inc.
Xuanhe Zhao @ProfZhaoMIT
24K Followers 2K Following Whitaker chair professor @MIT; scientist, educator, entrepreneur; #softmaterials #mechanics #health #sustainability; co-founder @SanaHeal @Sonologi @Magnendo.
Steve Cranford @CranfordMATTER
7K Followers 1K Following Editor-in-chief, Matter family, Cell Press. Dipping my toe back into X. Is it safe? Opinions my own. 🇨🇦
Nicholas A. Peppas @NPeppas
21K Followers 352 Following Cockrell Family Regents Chair, Chemical Engn Biomedical Engn, Dell Medical School, Pharmacy -- Institute of Biomaterials, Drug Delivery & Regenerative Medicine
Advanced Portfolio @AdvPortfolio
35K Followers 290 Following Accelerating Science. Empowering Scientists. Part of @WileyGlobal.
MIT MechE @MITMechE
26K Followers 447 Following Latest news and research from MIT's Department of Mechanical Engineering.
Jennifer A. Lewis @JenniferALewis1
6K Followers 727 Following Wyss Professor of Biologically Inspired Engineering @Harvard @wyssinstitute @hseas. Director, Harvard MRSEC - supported by @NSF & Entrepreneur
Markus J. Buehler @ProfBuehlerMIT
26K Followers 2K Following McAfee Professor of Engineering @MIT; Co-Founder & CTO at Unreasonable Labs; AI-Driven Scientific Discovery
80,000 Hours @80000Hours
32K Followers 515 Following You have 80,000 hours in your career. This makes it your best opportunity to have a positive impact on the world.
Carlos Santana @DotCSV
246K Followers 1K Following 🤖 Divulgador de Inteligencia Artificial (DotCSV) ✉️ Contacto comercial: [email protected] 📚 Enseño sobre IA en Youtube, Tiktok e Instagram
Max Welling @wellingmax
42K Followers 474 Following
ICML Conference @icmlconf
88K Followers 11 Following Int'l Conf on Machine Learning • This account not monitored, Contact: https://t.co/6saHKWVxR6 • LinkedIn: https://t.co/YlU0lrSVna
Hugo Larochelle @hugo_larochelle
125K Followers 649 Following Mila Scientific Director. Scientific Lead @adaption_ai, advisor @tiptreesystems & @PrizmalAi. Ex @Google DeepMind & Twitter Cortex. Father of 4.
Sergey Levine @svlevine
137K Followers 144 Following Associate Professor at UC Berkeley Co-founder, Physical Intelligence
Charles Sutton @RandomlyWalking
17K Followers 1K Following Research scientist @GoogleAI / Previously academic @InfAtEd / Deep learning to help people write code. / @[email protected] / ❤️s:🐱🐶☕️🍕
Durk Kingma @dpkingma
55K Followers 404 Following @AnthropicAI. Prev. @Google Brain/DeepMind, founding team @OpenAI. Computer scientist; inventor of the VAE, Adam optimizer, and other methods. ML PhD.
Cheng @zcbenz
8K Followers 105 Following maintainer of MLX @apple. creator of @electronjs. check https://t.co/ZDJujd4fAN for the open source things I built.
Andrew Yeung @andruyeung
78K Followers 821 Following hosting extraordinary people @meetfibe @theshortlistnyc | angel investor in 20+ companies | former @google @meta product lead
Paul Liang @pliang279
11K Followers 462 Following Assistant Professor MIT @medialab @MITEECS @nlp_mit || Foundations of self-evolving multisensory AI to enhance the human experience.
Hua Shen✨ @huashen218
3K Followers 1K Following ✨Assistant Professor of CS @NYU Shanghai, @NYU. Exploring bidirectional human–AI alignment, safety, and the dynamics of co-evolving systems. 🦋
Stefan Schubert @StefanFSchubert
53K Followers 2K Following I run The Update newsletter. Book: https://t.co/I5zN3WGe0p
Akanksha @akankshanc
2K Followers 875 Following Passionately in love with Science, Altruistic, Engineer, Amateur Astronomer & Critical thinker. Current Research focus: ▫️Mechanistic Interpretability▫️
Ruchir Jajoo @ruchirjajoo
11K Followers 38 Following I can think, I can write, I can wait Founder/CEO - Identity Labs
Mickey / Phase Transi... @MickeyXaman
54K Followers 11K Following Ex EM CEO & polymath—left/right brain, no center. Build econophysics models, write short stories. Phase transition here: let's learn/share. Not activist/advisor
Jack Lindsey @Jack_W_Lindsey
20K Followers 165 Following Neuroscience of AI brains @AnthropicAI. Previously neuroscience of real brains @cu_neurotheory.
Christina Wadsworth K... @ChristinaHartW
9K Followers 223 Following leading personalization & proactivity @OpenAI | previously @Meta, @Instagram
Vivek Kalyanarangan @vivekkquant
150 Followers 224 Following
Martin Tobias (Pre-Se... @MartinGTobias
60K Followers 10K Following Entrepreneur, Investor, girldad, cyclist, surfer, poker player. Pre-seed up to $500K. Chat with me https://t.co/96wsMImeiy. Get $$ https://t.co/d7utyst2XW
modest proposal @modestproposal1
126K Followers 799 Following I shall now therefore humbly propose my own thoughts, which I hope will not be liable to the least objection. https://t.co/NuC25VaSff
Patrick OShaughnessy @patrick_oshag
360K Followers 2K Following building @psumvc @colossusmag hosting @investlikebest
David Senra @FoundersPodcast
212K Followers 319 Following Learn from history's greatest entrepreneurs. Every week I read a biography of an entrepreneur and find ideas you can use in your work.
Jim O'Shaughnessy @jposhaughnessy
193K Followers 5K Following https://t.co/HGYG4UoYts 🎙 https://t.co/EgEFb3o46Q
Brent Beshore @BrentBeshore
131K Followers 1K Following @PermanentEquity // @CapitalCamp // @MainStSummit // Author of "The Messy Marketplace" // Former atheist who now follows Jesus
Christopher Bloomstra... @ChrisBloomstran
129K Followers 1K Following Nothing here is advice. Drop the phone and slowly back away from the Twitter. Read a 10-K. It’s a profession, not a business.
Gavin Baker @GavinSBaker
346K Followers 6K Following Managing Partner & CIO, @atreidesmgmt. Husband, @l3eckyy. No investment advice, views my own. https://t.co/pFe9KmNu9U
Brandon Beylo @marketplunger1
114K Followers 2K Following Writer | Investor | Listener || ~ There are no great businesses, only great bets. Nothing you read here is investment advice. Read 👉 https://t.co/rVOdJUnxQI






















