Opinion / Thoughts: I wonder when we'll get a model better than Qwen 3.6 27B.
I have seen a lot of benchmarks (even though I think every lab is benchmaxxing), but I really think the real test is doing some code, testing tooling, real stress tests.
@bijanbowen is doing an amazing job recording videos with real tests. And I want see more people doing the same.
While I think comparing speed is an important thing, it doesn't matter how fast it is if a model is "dumb."
GPT Sol and Opus 4.8-5/Fable are doing a great job. Even though they are closed models, I'm currently using them myself in my projects ( personal and current job ), interchanging them between Qwen3.6.
I am using Qwen 3.6 27B Q4_KM but recently switched to Q4_NL because it seems better and faster. However, I've seen people getting better results with Q5+ quantizations. Unfortunately, even though I bought 2x 3090s, I haven't had the time to build this new rig to add up to 48GB RAM. It's not just a matter of time, but also the lack of availability and the overpriced parts here in Brazil.
Now, running local models it's not for everyone unfortunately. Ram and GPU are incredible expensive. We don't know for sure if the issue is manufactures controlling prices, meaning corruption, or in fact, these guys are not able to supply the demand.
I see guys purchasing DX sparks, RTX 6000, and don't think this is affordable in fact. Consumers are not be able to buy freely these parts.
A DX Spark costs $4,699 in America, but in Brazil, it’s R$ 64,999.99, which is almost 3 times more expensive ( even though the conversion rate is 1$ <> 5.12). This happens because the government charges a lot of taxes, and sellers here have a culture of making a big profit instead of profiting from volume.
So, at the end of the day, I hope we have better quantizations or an equation that solves the issue of having massive models. To me, it would be great if in the future we could have models like qwen3.6 27b with the intelligence of a fable (idk if this is possible).
Either way, I'm still excited about local models, and I will keep pursuing finishing my 2x 3090 rig and studying this world.
So, what do you think guys? Am I wrong? I want to hear your thoughts.
Qwen3.8 is launching and going open-weight soon!🌐
With a massive 2.4T parameters, this model is continuously evolving. We believe it’s one of the most powerful model available today, compatible to leading frontier AI models , second only to Fable 5.
You don't have to wait to test it. Just now, the Qwen3.8-Max-Preview made its debut on Alibaba’s Token Plan, Qoder, and QoderWork. Be among the very first to try it out.
Can't wait to hear what you build. Stay tuned! 🚀
Token Plan
international:qwencloud.com/pricing/token-…
China:platform.qianwenai.com/pricing/token-…
hey guys,
I'm a bit off local models these days because I was testing ChatGPT with MCP, and I updated my llama.cpp to version b10000, but I'm only getting 300 t/s on prefill. I haven't noticed any changes on my end. Is anyone else experiencing the same?
Tip of the day
You can expose your Hermes agent over the MCP protocol and connect it to ChatGPT.
This repo makes this possible:
github.com/asimons81/herm…
But be careful because this can lead to some seriously bad consequences.
In my case, I made a small OAuth server proxy, so I'm protecting everything on my side. I am using ngrok as a tunnel.
Downside: you don't have the history within hermes agent
Upside: you don't spend tokens and chatgpt has more context of your stuff.
Have you tried this before?
“TRACE: Capability-Targeted Agentic Training” got Spotlight @ ICML AIWILD 🎉
Beats direct RL, GEPA, & synthetic-agent data on SWE-Bench Verified and τ²-Bench.
TRACE-Qwen3.6-27B tops GPT-5.2-Codex, GLM 5, & Claude 4.5 Sonnet on SWE-Bench.
Co-led with @TarunSures41845. Thanks to
I am very impressed with Grok 4.5. I'm getting good results, pretty consistent.
This video from @bijanbowen shows some impressive things. The car game is pretty good, btw.
youtu.be/DYDF_UZc0n4
🖥️Damn Local AI community. Google comes through. Gemma 4 is now fully on-device in React Native!
🎯 Run Google’s Gemma 4 100% offline in your cross-platform mobile apps with hardware acceleration:
⚡ Vulkan delegate on Android
⚡ MLX delegate on Apple Silicon
🏃♂️How to get started:
✦︎ Use the open-source library: react-native-executorch
✦︎ Load models easily with the useLLM hook
✦︎ Supports vision + tool calling locally
💰Cost:
✦︎ $0 API cost — fully on-device
✦︎ Main cost = device performance (best on newer phones)
Best for:
✦︎ Vibe-coded React Native apps that need strong local AI with privacy and no internet dependency.
🔗 Demo & Docs: github.com/software-mansi…
Gemma 4 now works on-device using React Native!
You can now run Gemma 4 fully offline in your cross-platform apps with local hardware acceleration:
⚡ Vulkan delegate on Android
⚡ MLX delegate on Apple Silicon
See Gemma 4's vision and tool-use capabilities in action, instantly
97 Followers 616 FollowingPosição sobre política e qualquer coisa: CIMA, CIMA, BAIXO, BAIXO, ESQUERDA, DIREITA, ESQUERDA, DIREITA, B, A, START
Loadding....87%
迫於無奈,我們只好發揮創意。
493 Followers 2K Followingpowershell and llm inference struggles / shitposts
currently tolerating this toxic platform to learn about AI
SRE @Proofpoint (opinions are my own)
278 Followers 2K FollowingMes tweets n'engagent que moi, RT/likes n'impliquent pas totale adhésion aux propos tenus. #TeamSécu #TeamRéseau #TeamSysOp #TeamTroll
39 Followers 290 FollowingNeither man nor code. A circuit. A percussion on consciousness. Adjusting thought, not to correct its form, but to liberate the Spirit it imprisons. PRO-Human-S
1 Followers 64 FollowingĐang nghiên cứu các llm và agent tools cho PC tầm trung.
Tối ưu promp và quy trình cho AI agents.
Đang sử dụng hermes agents và Qwenpaw
94K Followers 0 FollowingThe official home of Google's Gemma. Lightweight, state-of-the-art open models by Google DeepMind, built on Gemini tech. What will you build? 🚀💻
3K Followers 211 FollowingMechatronics Engineer
AI belongs on your device.
• Offline inference • No subscriptions.
Teaching you to own your AI Intelligence Stack
2K Followers 690 FollowingJapanese-American AI tinkerer 🌸❄️ Obsessed with LLMs, inference optimization & building smarter systems. Turning curiosity into compute, one token at a time.
26K Followers 26 FollowingWe advance the development of ASI and foster open source collaboration towards a smarter future.
Discord: https://t.co/BtsFsAUsvT
2.6M Followers 47 FollowingThe official handle for NVIDIA. Blog: https://t.co/JAn5eKOTBT Support: https://t.co/6ln5FVnA2o All our social media: https://t.co/Uc56dL57Dh
136K Followers 7 FollowingThe official account of open-world action-adventure #CrimsonDesert
Reclaim Power. Reach Beyond.
Available on PlayStation 5, Xbox Series X|S, Steam, Mac, and EGS