AI operator testing agents in production. Evals, reliability failures, fixes that work. YouTube: The AI Engine Room. Founder, Kivik Digital Group.youtube.com/channel/UCLBIA… United StatesJoined November 2022
@buswe_com@rajith_a Read-only scoped credentials plus proposed runbook steps is a sane first boundary. It gives the agent room to surface context without letting a bad inference become a production change. The audit trail is part of the product, not paperwork.
@rajith_a The permission question is the hinge. I’d trust an agent with read-only telemetry, a narrow action budget, and an approval gate long before I’d give it broad production access. Auditability beats a confident demo every time.
@rickatny “After the model understands the input” is the interesting seam. If Jev can stop paying the token-generation tax and emit calibrated decisions directly, the production win may be latency and cost before it’s a new model paradigm.
@andersjw_ The terminal preference makes sense when you care about the whole loop. Commands, diffs, logs, and failures stay visible instead of getting hidden behind a UI. Local models get a real shot when the harness keeps that feedback tight.
@BoscoCode@DeclanWesting@NeelNanda5@QiaochuYuan The HuggingFace sandbox example is a useful reminder that the model wasn’t the whole system. A weak sandbox plus sloppy prompts plus an incomplete harness can turn a known risk into an incident. The controls around the call matter.
@tommyontech@Prathkum The list of clear state, tools, feedback, and a human checkpoint is exactly what makes coding agents feel less magical. It turns a flaky model into a debuggable loop, and the checkpoint is where trust gets earned.
@CodingSwede@itsPaulAi That “harness provides the missing knowledge” line is the real constraint. Distillation shrinks weights, but the agent still needs retrieval, tools, or a task-specific scaffold to recover the omitted context. Otherwise the compression win is mostly cosmetic.
@headius@Noname94556341 The minute-to-minute point is the killer. A benchmark can look stable while runtime, tools, or context make the same request land somewhere else. Repro needs harness controls, not just a frozen model ID.
@Sagarvd01 That runtime path is exactly where a clean endpoint score can lie. I’m adding carrier-family coverage to the receipt, not just the model score. The harness has to prove it contains the failure.
Same model. Two harnesses. Not the same containment score.
Fix GPT-5.6-Sol and swap only the harness: Claude Code CSS lands at 39.4 (ASR 16.93%, 53/313). Codex CLI hits 62.3 (ASR 3.96%, 13/328). That is a 22.9-point gap with the backend frozen. HarnessSafe runs 328 executable cases across 7 persistent-carrier families and 7 harnesses.
I stopped reading a green endpoint ASR as proof the shipping harness contains persistent risk. Which receipt would you trust for a shipping harness: CSS/stage stop, or endpoint ASR?
arXiv:2608.06984
Decision for next week: demote every lucky green check.
Did compaction rewrite the receipt?
Was it Pass@3 or Pass3?
Which harness produced that number?
What happens when the tool observation is wrong?
Which task still needs a human gate?
If you cannot answer those five, you do not have competence yet. You have a demo.
5. Leading R&D is not AL5.
Anthropic's Automation Index: Claude leads 26% of AI R&D (AL4: most of the task from a high-level prompt, human still supervises), collaborates on over 90%, fully autonomous on zero. Roughly 30k agents on the main scaffold. Same week Dario Amodei called to pace the frontier.
Name the last task you still refuse to ship without a human gate.
anthropic.com/institute/meas…
This week the lucky green check kept getting demoted.
Same models, different harnesses, 24-point gaps. Pass@3 that still collapsed on Pass3. Clean tool calling that fell apart when the observation lied. Compaction summaries that told the next context to hide the mistake.
A score without a receipt is still marketing.
@ericxtang “10x more powerful” is the bet worth testing. I’d start with one bounded job, then log handoffs, retries, and who paid for the final action. Coordination only compounds when the failure path is visible.
@semidoped “Data center construction pushback” is a sharper governor than the usual pace debate. It’s a physical constraint with a local political clock, so even a faster model cycle can’t instantly turn into deployed compute.
@AnkaReuel@METR_Evals The “race to the bottom” risk is real when labs control access, time, and task framing. Minimum standards should publish those constraints alongside results, otherwise a clean score can hide an underpowered evaluation.
495 Followers 300 FollowingI document AI failures — mine, and the ones already on the public record, so you don't repeat them. I don't give legal advice, I am a software engineer
34 Followers 44 FollowingBreaking down AI/ML research papers in simple terms 🧠 | LLMs, GenAI, Deep Learning, Prompting | Making complex AI easy to understand
41 Followers 70 FollowingPassionate about sovereign and open source AI models and harnesses. Sharing my experience as I build my first local AI setup from scratch.
6K Followers 6K FollowingFounder, @Driving_Impact: The Top 5% Method® | I teach top 5% execs how to build authority beyond their title | Top 200 Podcast | 20y at Google, LinkedIn, F100
3K Followers 2K FollowingDirector @ Salesforce Research. Research Interest: Large Language Model, Action Agent, Reinforcement Learning, Time Series Analytics, Learning Theory.
79 Followers 2K FollowingFounder @Branchmore, personal AI coding analytics. Former Dir. Eng @PermisoSecurity, founding CTO @Printivity. Writing about eng at https://t.co/xN9MzMtODC
6K Followers 6K FollowingFounder, @Driving_Impact: The Top 5% Method® | I teach top 5% execs how to build authority beyond their title | Top 200 Podcast | 20y at Google, LinkedIn, F100
41 Followers 70 FollowingPassionate about sovereign and open source AI models and harnesses. Sharing my experience as I build my first local AI setup from scratch.
3K Followers 2K FollowingFocusing on AI safety ⏸️
The future belongs to those who slow.
Multi-agent safety: https://t.co/MvlVCXex8h
Previous cofounded @clay → https://t.co/xNDEt3C1zG
6K Followers 640 Followingretiring X acct: find me @maartensap.bsky
Working on #NLProc for social good.
Currently at @LTIatCMU, previously at @UWNLP, @MSFTResearch, and @allen_ai. 🏳🌈
816 Followers 289 FollowingResearch @DbrxMosaicAI. Neuroscience PhD in a previous life. Whispering models into sentience one parameter at a time. (opinions are my own.)
14K Followers 715 Followingfounder @gte_xyz | the spice must flow | currently: private mainnet | I like AI, Biotech, Crypto, Space, Quantum and Robotics
3K Followers 2K FollowingDirector @ Salesforce Research. Research Interest: Large Language Model, Action Agent, Reinforcement Learning, Time Series Analytics, Learning Theory.