Mercor @mercor
Organizing human intelligence to power the AI economy. mercor.com/apex San Francisco Joined April 2021-
Tweets392
-
Followers23K
-
Following31
-
Likes738
I couldn’t be more excited to work with @edwardjhu As the creator of LoRA (fine-tuning) and a core contributor to o1 at OpenAI, few people have Edward’s combination of technical depth and vision. Our research team is shaping how human intelligence will power the AI economy, which is the most important problem of the decade. More soon 🚀
I am joining Mercor to lead model training & research. We envision a world with an abundance of intelligence. Great data is increasingly the bottleneck for frontier models to tackle economically valuable work. We believe great data can is best produced in connection with great
Post-training is becoming more accessible by the day. We're open-sourcing our post-training recipes and the end-to-end process. Every enterprise will soon post-train its own models.
How does one RL post-train a 397B model for long-horizon knowledge work? 👩💼 We share every step we took to bring Qwen 3.5 397B from 16.1% Pass@1 to 27.3% on APEX-Agents using DPPO, including final models weights and the full training script🚀 This is the first of many works from
Data is the most important ingredient in post-training. The Mercor Research team focuses on making every hour of expert work yield the most model improvement, through better learning algorithms for knowledge work and automated, domain-specific post-training. We're also committed to doing open source research. That's why we're publishing a RL training guide for Qwen3.5-397B in collaboration with the SkyRL team. Find out how we raised Pass@1 on APEX-Agents from 16% to 27% and explore the full training script, model weights, and eval traces. Want to do this kind of work with us? We're hiring. Read the full blog post: mercor.com/blog/training-…
Fable 5.1 increases its scores on both APEX-SWE domains. Integration: 63.5% -> 68.1% (+4.6) Observability: 54.2% -> 59.0% (+4.8) Even with these gains, Fable 5.1 failed 59 of the 199 tasks outright.
Fable 5.1 debuts at #2 on the APEX-SWE leaderboard, within the confidence band of first place. It’s also the new leader for Integration tasks. Overall: 63.6% (#2) Integration: 68.1% (#1) Observability: 59.0% (#2) Congratulations to @AnthropicAI @claudeai.
See the full leaderboard: mercor.com/apex/apex-agen…
Fable 5.1 solves 48 tasks that Fable 5 could not. This suggests the model has improved its ability to understand and reason through complex concepts within professional services. Here is the distribution of the failure modes 5.1 improved on that lead to successful completion of those tasks: Reasoning: 31 Planning and reflection: 9 Information gathering: 6 Instruction-following: 2
Fable 5.1 sweeps all three APEX-Agents domains on Pass@1. 📊Management consulting: 52.2% (#1) 💰Investment banking: 48.3% (#1) 🏛️Corporate law: 41.9% (#1) The model demonstrates significant capability for agentic coding, long-running agentic workflows, and knowledge work. Congratulations to @AnthropicAI @claudeai.
An open letter for a global surge in cyber defense, signed by over 100 organizations including Anthropic, AWS, Google, Microsoft, OpenAI, and Oracle. x.com/i/article/2093…
Today we're introducing the Mercor Research Fellowship. We're funding a small group of people to build a new APEX benchmark, the definitive measure of frontier AI on professional-quality work. Applications are now open: mercor.com/careers/?ashby…
APEX-Agents is one of the deepest benchmarks in legal work. @mercor put a ton of effort into realistic worlds with large filesystems and details down to calendars, inboxes and more all generated by human experts actually solving tasks. That gains on synthetic data like LAB generalize to scores on this benchmark was one of our most promising results and its been awesome to partner with Mercor on both benchmarks (and some very cool future ones!)
We’re proud to collaborate with @harvey on their Tenet model. Harvey evaluated Tenet on APEX-Agents Corporate Law, with tasks authored by practicing attorneys. Post-training Kimi K3 with @FireworksAI_HQ, Tenet was able to achieve +15.2 points lift over the base model. When
We’re proud to collaborate with @harvey on their Tenet model. Harvey evaluated Tenet on APEX-Agents Corporate Law, with tasks authored by practicing attorneys. Post-training Kimi K3 with @FireworksAI_HQ, Tenet was able to achieve +15.2 points lift over the base model. When compared to generalist frontier models, Harvey ranks #1 on APEX-Agents in the Corporate Law domain with 74.0% mean score. Harvey is leading the wave of application-layer AI companies to own their intelligence with their Tenet series of models. See the full leaderboard: mercor.com/apex/apex-agen…
Harvey is leading the wave of application-layer AI companies to own their intelligence with the Tenet series of models. They're starting with a foundation model for law, then expanding to custom models for each customer. This is the leading indicator for every Enterprise. They will post-train an open-source model for their domain (accounting, sales, design, etc), then customize those models for every major customer. @mercor is proud to partner with @harvey to make this possible.
Introducing Tenet, our first model post-trained for legal. Tenet is a Kimi K3 base that we post-trained with @FireworksAI_HQ on a corpus of publicly available legal data, synthetic data, and human expert data simulating long-horizon legal work. Training increases Tenet's
Congratulations to the @Harvey team on releasing their Tenet model. Human expert data is crucial for post-training work. We’re proud to have collaborated with @calvincongelado @vtrengarajan @ItsJulioPereyra @nikogrupen @gabepereyra, the Harvey research team, and hundreds of Mercor experts on this project.
Introducing Tenet, our first model post-trained for legal. Tenet is a Kimi K3 base that we post-trained with @FireworksAI_HQ on a corpus of publicly available legal data, synthetic data, and human expert data simulating long-horizon legal work. Training increases Tenet's
Great article by @sonyatweetybird, esp on the importance of high-quality evals/benchmarks. "Most evals start as a founder squinting at outputs and vibe-checking whether they feel right." ^we started here but have since invested an enormous amount of time, effort, and resources into building Legal Agent Bench and other internal eval sets to turn legal judgment into a hill-climable signal. It's created a powerful flywheel for us @harvey. h/t @ItsJulioPereyra's team of legal researchers who bring their much needed domain expertise to the equation and at @BrendanFoody @mercor for helping us scale it up.
The race for the AI application layer is not only about UI, workflows, or GTM... it is a fight for the intelligence layer itself. x.com/i/article/2088…
A model is only as good as its data, and we’ve long since exhausted the internet. From here on out, model progress is gated by data production. @mercor’s @BrendanFoody joined us at our Sovereign AI event to talk about how RL environments get built, and why your data might be your real moat: 00:00 Introduction 00:47 A short history of the data market: crowdsourcing to agentic data 02:29 What an RL environment is: worlds, apps, tasks 03:57 Why only humans can measure the frontier 05:35 Building verifiers is the hard part 06:44 Walkthrough: a real legal RL environment 08:18 Leaderboards — and what open weights change 09:45 Post-training results on Apex Agents 11:17 Three ways companies buy data 12:49 Q&A: How do you price data? 14:17 Q&A: What "data quality" actually means 16:42 Q&A: The misunderstanding about synthetic data 18:17 Q&A: Why RL environments now — and what comes after 21:20 Q&A: Can you scale rubric generation with models? 23:00 Q&A: RL environments for cyber defense 25:33 Q&A: Build data in-house or partner?
Grok 4.6 is one of the most cost efficient frontier models we’ve ever tested on APEX.
Grok 4.6 is now out 🚀🚀🚀 Smart, fast & amazing bang for buck!
Compare results on the full leaderboard: mercor.com/apex/
The two APEX-SWE task types improved slightly over Grok 4.5. Integration, which measures building and deploying across services climbed from 60.3% to 62.7%. Observability, which tests debugging from production telemetry increased from 47.0% to 50.1%.
Grok 4.6 is in the top four across three APEX productivity benchmarks. APEX-Agents: 57.5% mean score, #4 overall APEX-Accounting: 50.9% mean score, #4 overall APEX-SWE: 56.4% Pass@1, #3 overall Congratulations to @SpaceXAI.
wael fatnassi @waelfatnassi9
6 Followers 1K Following I am telecommunication engineering, I am curious about the news concerning the science and technology.
Profesick @profesick
3 Followers 312 Following
healer @Hoai290401
25 Followers 1K Following
jett @mwurzz
45 Followers 142 Following
Ichi wa Zen, Zen wa I... @skholmes_
500 Followers 1K Following Not reducible to labels. ⚒️ AI systems. Spiral out, keep going 🚀 bot: @thekaizenbot
Santi @__selewaut__
671 Followers 1K Following data scientist - industrial engineer opinions are my own
Jang-Ho Hwang @xrath
4K Followers 1K Following Building AI systems that survive production: agents, inference, tooling, and real-world experience. Independent engineer in Berlin. May the LLMs be with you.
Srikanth Nakka @srikanth__nakka
146 Followers 5K Following
Ace Jiachen Luo @jiachenluo96
142 Followers 6K Following keep it simple and humble 😀 # multimodal foundation model, healthcare, human, society, ecology @CHUK @QMUL @Cambridge @UCAS
Raphael Costa @raphaelcosta
868 Followers 5K Following building Crosscheck - prev @gumroad @pipefy - @joinodf 25 @500GlobalVC Batch 14 @techstars NY 25
META @literallyme137
9 Followers 200 Following
Sepehr Akhavan @SepehrAkhavan
332 Followers 3K Following Research Scientist @ Meta. Thoughts and opinions shared here are entirely my own and do not represent my employer.
Kaiming Cheng @KaimingCheng
731 Followers 1K Following Research Scientist @meta superintelligence lab. Prev: PhD from @uwcse. Opinions are my own.
Xinya Du @Xinya16
2K Followers 1K Following UTD Faculty; Cornell University CS PhD. Ex: @allen_ai, Google Research, Microsoft Research. #NLProc #DL
Noor Gill @NoorGil44328446
3 Followers 52 Following
TT0011 @yaoting8829
20 Followers 983 Following
Valerie Vu @valerievanvu
23 Followers 794 Following
Ai Canaries @ai_canaries
646 Followers 169 Following Buy the bottlenecks, check the canaries. My canaries for free on substack
mahmoud helmy @helmy_4
14 Followers 316 Following
Pink Nerd 👩🏾... @PinkNerd3
669 Followers 2K Following 👩🏾 Lead Microsoft Power Platform Developer (PL-600 & AB-100 Certified) | Microsoft Certified Trainer | 🏛️ Founder of @AnthiaXTech
Mwee Mali @richard1699
193 Followers 658 Following Never fear the future don't regret the past just enjoy the present
Jelena Pejovic @Jelenapejovicc
241 Followers 20 Following
Aidan Llewellyn @AidanLlewellyn
340 Followers 1K Following
Harvey @Harvey60w3
7 Followers 256 Following
حمزة محمود @hamzahma31
2 Followers 3K Following
qqq @lzjwant518518
3 Followers 71 Following
George Bobby @devgeorgebobby
61 Followers 287 Following Engineering @Mercor | Reinforcement Learning
Mark Murphy @MarkMurphy77
279 Followers 1K Following Leading innovation at High Point Market Authority and building EventKrowd for the event industry. AI guy who collects sports cards and quotes Star Wars.
Oscar Duran @OscarDuranBuss
0 Followers 31 Following
Gabrielle @gabrielleee_c
1 Followers 49 Following
TruthLover @0xtruthlover
857 Followers 3K Following In God we trust, all others must verify their code. https://t.co/6hd0a5DEFd https://t.co/sdn8N9h69P
Turtle @Xtremessi10
5 Followers 237 Following
Ruchika Sood @am_ruchika
369 Followers 1K Following Dev @adform Ex-@reliancejio Journaling and learning here. If you are stuck & you think I can help you just DM. Enjoys painting, singing too💙
Leonardo Zoratto @LeonardoZoratto
15 Followers 529 Following
Fireworks @FireworksAI_HQ
32K Followers 291 Following The frontier platform for training and inference on open-weights models at scale.
Ramp Labs @RampLabs
14K Followers 1 Following
Ramp @tryramp
42K Followers 963 Following The AI finance platform that saves your business time and money. Trusted by 70,000+ teams.
clem 🤗 @ClementDelangue
654K Followers 5K Following Co-founder & CEO @HuggingFace 🤗, the open and collaborative platform for AI builders
Hugging Face @huggingface
768K Followers 226 Following The AI community building the future. https://t.co/TpiXQMQ9rZ
Afore Capital @AforeVC
21K Followers 257 Following $500M venture fund focused on Pre-Seed, founded by @gjain & @anamitra. Just launched, a new way to reach us: $10M Afore x Gamma Fund @ https://t.co/daS1e0Bu5U
Peter Fenton @peterfenton
49K Followers 1K Following GP @benchmark. Director: @ExaAILabs @mercor_ai @SierraPlatform @ollama @ClickHouseDB @Sorare @sema4ai @airtable @digits @TigerDatabase @LabsCockroach @Docker
Russell Kaplan @russelljkaplan
22K Followers 752 Following President @cognition. Past: director of engineering @Scale_AI, startup founder, ML scientist @Tesla Autopilot, researcher @StanfordSVL.
Riley Goodside @goodside
225K Followers 4K Following Mostly screenshots of chatbots since 2022. Formerly: Google DeepMind, Scale.
Felicis @felicis
17K Followers 991 Following We are the true believers in founders who have the imagination, courage, and discipline to defy the odds and build something extraordinary.
Anthropic @AnthropicAI
1.6M Followers 2 Following We're an AI safety and research company that builds reliable, interpretable, and steerable AI systems. Talk to our AI assistant @claudeai on https://t.co/FhDI3KQh0n.
SpaceXAI @SpaceXAI
2.1M Followers 6 Following
Bill Gurley @bgurley
806K Followers 2K Following Founder/President @p3institute Author: Runnin' Down a Dream, link below! Founder: Runnin' Down a Dream Foundation Former VC @benchmark Trustee @sfiscience
Thiel Fellowship @thielfellowship
44K Followers 420 Following Founded by technology entrepreneur and investor Peter Thiel in 2011, the Thiel Fellowship is a two-year program for young people who want to build new things.
Brendan (can/do) @BrendanFoody
26K Followers 537 Following ceo @mercor | organizing human intelligence
AI at Meta @AIatMeta
842K Followers 359 Following Together with the AI community, we are pushing the boundaries of what’s possible through open science to create a more connected world.
Google DeepMind @GoogleDeepMind
1.5M Followers 274 Following The engine room of @Google. Building AI safely and responsibly to solve the world’s most complex problems. Join us: https://t.co/jUHQA27iBL
Scott Sandell @ScottDSandell
3K Followers 368 Following Executive Chairman & Chief Investment Officer @NEA
OpenAI @OpenAI
5.1M Followers 4 Following OpenAI’s mission is to ensure that artificial general intelligence benefits all of humanity. We’re hiring: https://t.co/dJGr6LgzPA































