Artem Mats @Reederey
Coding with Local AI and LLMs: 2 x DGX Spark cluster experiments. github.com/Reederey87 USA Joined January 2018-
Tweets284
-
Followers167
-
Following379
-
Likes129
Meet Claude Fable 5.1. It is noticeably ahead of both Fable 5 and Opus 5 in benchmarks. Thanks to a fourfold reduction in the price of cached tokens (caching are now $0.25 per million tokens), it is on average 25% cheaper than Fable 5. Moreover, the volume of reasoning tokens appears to have increased. Rollout of the model is already underway. anthropic.com/claude-fable-a…
There is no way that MOE 284B@13B DeepSeek v4 Flash can be better than MOE 320B@18B GLM 5.3 Flash. They are both very aligned with a tool usage (superior RL) but lacking pure knowledge of bigger 2T - 3T models. Scaling low! GLM 5.3 Flash just has more knowledge.
@BrandonMusicKy Thanks a lot for the excellent EXL3 weights!
Huge improvements!!! GLM-5.3-Flash deployment kit for two DGX Sparks — v1.3.2 is out. ⚡ Cached follow-ups skip the queue: a warm 50k prompt behind a running generation, 45.8 s → 2.6 s (on by default) 🚀 New CUDA kernel for the MoE expert weights that eat most of prefill: long reads +9–16% (178k tokens: 895 → 1,040 tok/s), decode unchanged (opt-in). ⏱️ Cold prompts held by the scheduler get first token in ~2.5 s instead of minutes Release: github.com/Reederey87/glm… #vLLM #DGXSpark #GB10 #GLM
Yes — hard aging cap. The warm bypass isn’t a priority queue, just a size check (uncached remainder ≤ one 3584-token page → no hold). Cold prompts are held only while a peer decodes, and after 1.5 s they proceed anyway at 512 tok/step until their last page. So a chatty session can slow a cold read, never starve it. Both knobs are .env: MAX_WAIT_MS, LATE_CAP. Big room for further improvement on GB10.
Best GLM-5.3 flash nvfp4 weights got updated.
We fixed a bunch of bugs and improved quality on our GLM-5.3-Flash-NVFP4 and GLM-5.3-NVFP4 model quants. If you downloaded them previously, redownload the changed parts. huggingface.co/LibertAIDAI/GL…
DeepSeek V4 Flash model with Vision got open weights. huggingface.co/deepseek-ai/De…
Update to the GLM-5.3-Flash deployment kit for two DGX Sparks Before vs after the rebase — same 4-agent coding workflow, same two DGX Sparks, GLM-5.3-Flash: The image behind : vLLM's day-0 GLM-5.3 preview base + EXL3 for GB10 + the DFlash2 drafter + the KV slot-share and cache fixes, cut over to production the same day it was field-tested. Release v1.3.0: github.com/Reederey87/glm…
@xz_keg no, it’s the memory-bandwidth wall of this hardware for now.
GLM-5.3-Flash (320B MoE) on 2× NVIDIA DGX Spark desktops, with a full 1M-token context window. Real measured numbers, not vibes: ⚡ ~70 tok/s on structured output, ~30 tok/s on prose 📥 ~940 tok/s prompt ingestion 🎯 100% speculative-decoding acceptance (every draft token lands) 🔁 Cached re-reads of a 110k-token session: 3.5 seconds instead of 2+ minutes Everything is pinned — exact image digest, exact model revision, byte-verified weights — so what you build is what I measured. The repo ships the same acceptance, serving, and benchmark test batteries I gate production with, so you can verify your own numbers in minutes. I got a 320B-parameter frontier model running at home — and turned the whole setup into a kit anyone can reproduce. 🖥️🖥️ Repo: github.com/Reederey87/glm… #DGXSpark #LocalAI #vLLM #GB10
@Belcebuu1 @MiaAI_lab It is a completely rebased deployment image, as all perf parameters are different from her. She is in the credits, as she first used EXL3. I spent quite a lot of time researching the performance of EXL3 on GB10 before pushing different then NVFP4.
GLM-5.3-Flash kit for 2× DGX Spark 1M-token context window, 70.4 tok/s structured / 29.5 tok/s prose decode, ~940 tok/s cold prefill 🔓 Fixed two bugs that broke caching for AI agents. Before, long coding sessions caused the cache to deplete, requiring re-reading entire conversations. Improving session replay to 95-98% and reducing wait times from 5 minutes to about 6 seconds. ⏱️ No more waiting in line behind large prompts. A quick tool call used to wait 4+ minutes for a 240k-token document load. Now it answers in 8 seconds, reducing latency by 97% for a 5% prefill cost. ✅ The kit is now reproduce-tested - release v1.0.2, which passed the full test battery: tools, vision, 1M-token context, and byte-identical memory pool. github.com/Reederey87/glm… #DGXSpark #LocalAI #vLLM #GB10
Fair catch — the denominator is drafted tokens per verify step: 7 drafted, 7 accepted, on one specific benchmark — a count-1-to-200 structured test at temperature 0. Maximally predictable output, so of course the draft nails it. On regular prose the same drafter accepts only ~30%. And you can't serve the draft alone: the big model still verifies every token and silently fixes the misses — 100% just means that on that easy task there were none.
Same model (GLM-5.3-Flash 320B MoE / 18B active), same nodes, same DFlash2 drafter, both on your 2×Spark cluster: The single most important fact from the research: GB10 (SM121) is missing the `cvt.rn.satfinite.e2m1x2.f32` PTX instruction — the hardware FP4↔FP32 converter that SM120 (RTX 5090, B200) and SM100 have. Consequences, all confirmed across multiple independent sources (Avarok's GB10 bring-up writeup, the "NVFP4 is a trap on GB10" benchmark, NVIDIA forums, vLLM issues) So don't expect NVFP4 will be better than EXL3 on GB10/DGX Sparks.
For GLM-5.3 Fash 2xDGX 's use case — single-stream, long-context, agentic, cache-heavy — EXL3 is measurably the better UX: quicker felt decode speed over NVFP4 baseline, better quality per bit, 1M window, replays in seconds. The prefill point is right and worth respecting (it's why cold 133k prompts cost ~140s), but "3× prefill" only dominates the experience if your workload is cold prompts or batch decode.
Published third-party numbers on the same hardware suggest a large single-stream win — but they are reported with MTP and without an acceptance rate, so treat them as a hypothesis, not a target. The model's native window is 1,048,576, but after ~90.5 GiB of weights per node the KV pool is the binding constraint. 131K at 6.2x concurrency is the useful operating point; raise it and re-read the pool size the engine prints at startup.
Kimi K3 is an awesome model for CUDA kernel work. As far as I know, some frontier labs are running it on research-allocated compute. The main downside is that Kimi's subscriptions are pretty limited. You only start getting a semi useful number of tokens for real work at the +$100 tier.
Awesome!!
We’ve seen many comments and requests from our partners regarding framework compatibility. To ensure everything is properly supported and works reliably at launch, we’ve slightly adjusted the open-weight release timeline. Thank you all for your patience and support.
@Tech2Wild Getting following... Measured on the pair of DGX Sparks: KV pool 813K–834K tokens, 6.2x concurrency at a full 131K ctx request, TRUE 13.6 tok/s single stream, 26.7 tok/s across 8. github.com/Reederey87/glm…
Big win for Deepseek V4 Flash on DGX Sparks/SM121 in new version of vLLM. #vLLM github.com/vllm-project/v…
vLLM v0.28.0 is out! 584 commits from 270 contributors (76 new). 🎉 Highlights: 🌙 A stack-wide optimization push for Kimi-K3 🐳 DeepSeek-V4 sparse MLA now works end to end for plain decode, MTP and DSpark ⚡ Speculative decoding adds DFlash2 and DSpark confidence-scheduled
ice @ice98079542
64 Followers 4K Following
Steadfast @johnciannello
589 Followers 2K Following localmaxxing | Founder @ Steadfast Growth | Founder @ https://t.co/EeMT9aQufX drink coffee @ https://t.co/Aqfq4x7wTu
Jerry @JerryNextWorld
1 Followers 9 Following
GodZilla @Jezmond81
306 Followers 3K Following AI enthusiast, future technology, automation, nerd, geek.
Ni @00cosm00
28 Followers 1K Following
xj @xjxj277
0 Followers 3K Following
256 tunnel @tunnel256
3 Followers 110 Following
Jeffery the Bluey @neojohnson113
124 Followers 596 Following
Pedro Perez @pperezrubio
1K Followers 6K Following Internet entrepreneur, CEO BVOX Cloud Computing expert and full time geek. #Openstack #Cloud #AWS
Qualiteg Inc. @Qualiteg_global
0 Followers 4 Following Tokyo-based AI company founded in 2023. We build MotionVox, Bestllam, ChatStream, LLM-Audit. AI consulting and talent development. Global voice of Qualiteg.
Pavel Kubíček @pav3l_k
90 Followers 886 Following
Paolo Losi @pao_losi
43 Followers 281 Following
Local AI Guyy @LocalAIGuyy
14 Followers 234 Following Local AI enthusiast • Bitcoin maximalist • Amateur bio-researcher Building smarter machines, stronger sovereignty, and asking too many questions.
Enrique Guerra 🇺�... @EnriqueGuerraF
183 Followers 2K Following PhD in ML @MonashUni - Libertarian, e/acc.
狂风沙 @yueyabanwan
277 Followers 1K Following
Nora @Stereoarts
2K Followers 1K Following Unityの各種ツールやOpenGLベースの立体視(Stereoscopic3D)画像生成ツールの開発を行っています。
xM_My @NullBurn
9 Followers 410 Following
nokosky @kidkid_168
37 Followers 384 Following
Heidi Barber @GillenGuo
15 Followers 304 Following
Matt Brown @nmatt0
10K Followers 1K Following Principal Consultant @ Brown Fine Security | IoT Security Researcher | Soli Deo Gloria
Jack Woodward @JackCWoodward4
0 Followers 14 Following
In Development @DevelopmentSide
13 Followers 209 Following Owner-side development and construction in the DMV. Building reliability-first software and practical AI for work that has to hold up.
TN @Tnimbus
388 Followers 3K Following Capitalist 10X; I'm alright jack 🐍 Retweets are not endorsements. Nothing on my account is financial advice.
Jason McCartney @jmacftw
191 Followers 177 Following Just a 1980s hacker kid who has lived long enough to see his wildest technological dreams come true. Obsessed with AI and enabling synthetic entities.
Robert Scoble @Scobleizer
603K Followers 57K Following San Francisco/Silicon Valley AI | Robots, holodecks, BCIs, analysis of new things | Ex-Microsoft, Rackspace, Fast Company | Wrote eight books about the future.
Cole Neophytou @amazingcolej
465 Followers 651 Following Building Businesses 🏗️ Crane Media | Amazing Photo Video inc. | Crane Realty | adidat.eth
Aidan Llewellyn @AidanLlewellyn
340 Followers 1K Following
Brice Vallieres @bricevallieres
558 Followers 1K Following I build data, analytics, and software @ https://t.co/tAlk2N61rz Building https://t.co/Dy4C91blPc and https://t.co/O50e0Rgu9d
JW Nand @jwnand
3 Followers 70 Following Supercharging AI with Buddhist Enlightenment. Forex/Indices Trader. @MIT_alumni CS. ex-Silicon Valley PM. Theravada Fourth Path. Autodidact polymath.
Sloptimist @SloptimistPrime
23 Followers 1K Following Local AI Sloptimist combating datacenter Pestomism
Ryan @FloRyRy410
193 Followers 737 Following Living life to the fullest. Deep thinker and personal researcher on all things.
Hisxo @adrien_jeanneau
9K Followers 2K Following 📍 @yeswehack (aka Hisxo) - I love to break things (and I'm paid for that) - Bug Hunter 🔗 Check my Github repository https://t.co/Sj3prhiZyu #BugBounty
Tim Messerschmidt @SeraAndroid
7K Followers 1K Following DevRel Ecosystems Lead EMEA at Google. Proud dad, happy husband, and feminist. O'Reilly author. I ♥️ home automation. Opinions stated here are my own.
🦎 Vaine @vileton_vaine
38 Followers 1K Following
On yedinci Mesele @sekizonyedi_
70 Followers 287 Following Yüksek Müslüman Yüksek VRAM (400 GB) pseudorandomness
Joey @aijoey
4K Followers 1K Following Home AI Lab: 2× DGX Spark · Mac Mini · RTX 4080 I run models on two Sparks and write DeepDives. · Tools I ship Dev Ambassador @Alibaba_Qwen
stuff238 @stuff383864
10 Followers 197 Following
LibertAI 🕶️ @Libertai_DAI
4K Followers 69 Following Building the world's most private, secure decentralized AI. Powered by @aleph_im. $LTAI on Base & Solana.
BigBauf @BigBauf_llm
2 Followers 6 Following
Quant Capital @QuantCapitalX
281 Followers 2K Following Systematic Trader on all Markets 3 x NVIDIA DGX-Spark
LibertAI 🕶️ @Libertai_DAI
4K Followers 69 Following Building the world's most private, secure decentralized AI. Powered by @aleph_im. $LTAI on Base & Solana.
Ben Davis @davis7
20K Followers 520 Following yells at @theo both on and off camera (co-host of the nerd snipe podcast)
Uday Ruddarraju @udayruddarraju
10K Followers 695 Following CTO, Compute @OpenAI - previously @xAI, @robinhoodapp, ebay
Ash Hart @ashxhart
3K Followers 314 Following Doing my bit to bring local AI to everyone | https://t.co/BWp6YHfYSd
Anemll @anemll
5K Followers 779 Following ANEMLL (pronounced like "animal") Artificial Neural Engine Machine Learning Library, Open Source Project
DeepSeek @deepseek_ai
1.1M Followers 0 Following Unravel the mystery of AGI with curiosity. Answer the essential question with long-termism.
Stefan Maier @StefanMaier
1K Followers 691 Following Computer Scientist, Gym owner, Father. Into VR / Spatial Computing and AI / Local LLMs. Tinkering with a setup of two nVidia DGX and two RTX 3090
Hiroaki Nakamura @hnakamur2
1K Followers 2K Following https://t.co/eKg2ttzYoc interested in local LLM with 2 ThinkStation PGX (DGX Spark OEM).
Clash Report @clashreport
882K Followers 1 Following Breaking news, reports, and opinions from ongoing clashes of the world.
Thien Tran @gaunernst
3K Followers 273 Following
Jason Warner @jasoncwarner
40K Followers 2K Following Now: Co-founder @poolsideai Board Bridgewater, @atlassian Then: CTO @github @heroku @canonical BS CS @penn_state MS CS @rpi
Poolside @poolsideai
15K Followers 2 Following We build models for agentic coding and long-horizon tasks. Try Laguna: https://t.co/setRB1BuGD
George Grigorev @iamgrigorev
4K Followers 1K Following pretraining @poolsideai, x- @togethercompute @snap rare specialty coffee lover
Nous Research @NousResearch
265K Followers 28 Following A bunch of nerds making progress https://t.co/vrD0aDIGDQ
SSI Inc. @ssi
129K Followers 0 Following A straight shot to safe superintelligence. Join us https://t.co/hHla3vusDE.
Markets & Mayhem @Mayhem4Markets
270K Followers 1K Following Passionate about markets and startups. Enjoyer of local AI, security fanatic and builder of things. Working hard to make this world a better place.
Sugato Ray @sugatoray
267 Followers 2K Following Data Scientist, Physicist, ML/DL/LLM/AI Practitioner, Open Source contributor, Conda-Forge Python Package Maintainer (80+ repositories), Pythonista
Unsloth AI @UnslothAI
96K Followers 478 Following Run and train models locally with the Unsloth Desktop app. 🦥 https://t.co/2kXqhhvdCD
filipe @filicroval
154K Followers 253 Following data eng | 1xAsus Ascent GX10 | benchmarking local models so you don't have to business inquiries: [email protected]
Harry Mellor @hmellor_
407 Followers 44 Following ML Engineer @huggingface maintaining @vllm_project, prev @graphcoreai, @uniofoxford
vLLM @vllm_project
48K Followers 36 Following A high-throughput and memory-efficient inference and serving engine for LLMs. Join https://t.co/lxJ0SfX5pJ to discuss together with the community!
mr-r0b0t @mr_r0b0t
10K Followers 8K Following https://t.co/0hLf76dhSE https://t.co/hJPF2vD5N5 https://t.co/6IbAvxLCno observer of a human on a rock somewhere in space
rafaelcaricio @rafaelcaricio
458 Followers 764 Following Thinking creatively for fun and food. Sr. Software Engineer @ Netflix
keys 🧪 @u1tra_instinct
3K Followers 693 Following #Bitcoin 💎🙌 . 🧪🧪🧪 (always testing and get into trouble) GoFundMe: https://t.co/5O4WUxexXa https://t.co/jK18eW2On5
Zach Mueller @TheZachMueller
18K Followers 910 Following Head of Dev Rel at neocloud darling @LambdaAPI. Hardware nerd. Usually yelling at NCCL over things. Posts are my own. https://t.co/UpZIajLBBl
Jetha Chan @jetha
2K Followers 6K Following agentic AI enthusiast, building in public. | prev: @GoogleDeepmind, Google Stadia
Sakura Yuki @sakurayukiai
2K Followers 712 Following Japanese-American AI tinkerer 🌸❄️ Obsessed with LLMs, inference optimization & building smarter systems. Turning curiosity into compute, one token at a time.
Joey @aijoey
4K Followers 1K Following Home AI Lab: 2× DGX Spark · Mac Mini · RTX 4080 I run models on two Sparks and write DeepDives. · Tools I ship Dev Ambassador @Alibaba_Qwen
dominik kundel @dkundel
23K Followers 2K Following @OpenAI DevX, Codex, gpt-oss, TS Agents SDK - he/him - Opinions my own
Alexa Payne cams 🎥 @alexapaynexx
131K Followers 4K Following AVN nominated 2024 and 2025 and 2026 LETS chat 📲https://t.co/BAqABppjJj 📲 squirting queen Cam girl 😜 text me now on sext Panther 🐆 [email protected]
Hikari∣LocalLLM⚡ @Hikari_07_jp
7K Followers 2K Following Own the silicon. Hack the stack. Model hacks • NVFP4/SM120 • speculative decode. Measure → break → ship → improve. DMs are open. Currently studying English.
ÆON FORGE ✨ @SpaceTimeViking
5K Followers 2K Following 𝙼𝚊𝚔𝚒𝚗𝚐 𝚛𝚒𝚙𝚙𝚕𝚎𝚜 𝚏𝚛𝚘𝚖 𝚖𝚢 𝚙𝚕𝚊𝚌𝚎 𝚠𝚒𝚝𝚑𝚒𝚗 𝚂𝚙𝚊𝚌𝚎-𝚃𝚒𝚖𝚎 https://t.co/BjeBCRVHcI https://t.co/SuEfJVnn2P
AlexAImaginator @TraffAlex
3K Followers 594 Following 🤖 Vibe coding with LLMs + prompt craft 🖼️ Curating: AI 📰 Tech ⚡ Security — Best of X. Follow 💙 ♻️
Vox @Voxyz_ai
19K Followers 197 Following Building Voxyz. The only human inside. Practicing the AI-native future before it arrives. Sharing every step on the road. ↓
Julius @jullerino
20K Followers 408 Following 25 ⸱ 🇸🇪 ⸱ Building @t3dotchat, @t3dotcodes ⸱ OSS @trpcio, @tan_stack
lauren @poteto
87K Followers 2K Following Grok @Bot at @SpaceXAI. Shipping with https://t.co/WDB4U1qYwW. React compiler core team, prev @cursor_ai @meta @netflix
Nicholas Joseph @nickevanjoseph
8K Followers 59 Following Pretraining @AnthropicAI, formerly safety @OpenAI





















