Built by model builders, for model builders.
Humanlaya engineers expert-level data and verifiable rewards for frontier AI labs.humanlaya.comJoined March 2026
Why it matters + resources 🔗
If we want agents in finance/law/medicine/engineering, we need evaluation that matches real workflows + real cost of mistakes.
Dollars Matter, Scores Flatter. 💰
Can LLM agents really do high-valued experts work end-to-end — not just “solve a question”?
Introducing $OneMillion-Bench benchmarks real expert work, priced in dollars 💵
What moves the needle 🔧 Tools and evaluation design matter a lot.
• Official scaffolds > OpenRouter > no search (consistent gains)
• Rankings are stable across different judge models (robustness check)
• We also stress-test temporal sensitivity and test-time scaling behavior.
How Does We Grade? ✅ ❌
We only count value when the output is deliverable, not “partially correct.”
• 15–35 expert rubrics per task + negative rubrics for hallucination/unsafe/non-compliance
• Asymmetric weights (-20 to +10) to reduce reward hacking
From high-economic-value tasks to verifiable rewards, we help frontier teams push model capability beyond benchmark optimization and toward real-world cognitive performance.
We design expert-level tasks, build rigorous verification systems, and turn high-density human knowledge into scalable training signals for advanced models.
93 Followers 241 FollowingOSS AI Investor | Public OSS Investment Scorecard V1.2 🧪 | Singapore Top VC EIR | Helping builders evaluate open source AI & VCs find the next 10x | Taipei | g
158 Followers 474 Followinghttps://t.co/QxL5EK1BUr | AI Developer. | AI4AI | RSI
All views expressed do not represent any individual or the company’s position.
172 Followers 268 FollowingBuilding ATG, Research @Columbia @FieldsInstitute | AI, Quant Research | Accelerating math reasoner and verifiable financial intelligence. Views are my own.
74 Followers 236 FollowingPersonal market notes. I'm writing here for myself and to track my own ideas. I am not giving any investment advice. Too retarded to fail.