jupiter @jupiter186
Joined February 2009-
Tweets2K
-
Followers47
-
Following235
-
Likes278
We use GLM-5.3-Flash build a dream kitchen. A 3D world built in Blender. This is not a generated video.
这个台词水准有点高,三四层楼那么高。👍
Inspired by @percyliang's CS336, I built a static performance model for LLM inference — no dynamic batching/chunking, just clean analytical bounds. Pick model × GPU × batch × seq length × parallelism(DP/TP/EP/PP) → VRAM check + TTFT + TPOT + prefill/decode breakdown + throughput vs batch size. Covers 5 KV cache variants (GQA/MLA/SSM/sliding window/linear attention), speculative decoding, multiple quant precisions. Calibrated against ~100 public benchmarks (TRT-LLM, Splitwise, MLPerf, Koyeb…). Demo: …inference-calculator-delta.vercel.app Repo: github.com/pochenai/llm-i…
Top open models by use case 👇 chat - Kimi K3, Qwen3.8 2.4T A95B reasoning - Kimi K3, DeepSeek V4 Pro coding agents - Kimi K3, DeepSeek V4 Flash, GLM 5.2 small + fast - Qwen3.8 27B, gpt-oss 20B mid-size general purpose - DeepSeek V4 Flash, MiniMax M3 function calling - DeepSeek V4 Flash, GLM 5.2 vision - Qwen3.8 2.4T A95B, Kimi K3, MiniMax M3 speech-to-text - NVIDIA Nemotron 3.5 ASR, Whisper Large V3
Hotchips 2026 干货和私货分享 - 内存 - 知乎 zhuanlan.zhihu.com/p/207561757526…
From Zero to GPU: A Guide to Building and Scaling Production-Ready CUDA Kernels huggingface.co/blog/kernel-bu…
我小侄子大学三本毕业了,看了四年剧,天天吃外卖,问我有没有赚钱好的路子? 我说那肯定有,体面三件套。考研、考公、考编,他摇了摇头说太卷了,读书太累考不了。 我说那就牛马三件套。进厂拉货、做销售,他说想自由点的。 我说那就铁人三项。快递、外卖、网约车,他说有没有安稳点的。 我说那就只有吉祥三宝:保安、保洁、保姆。他说来钱太慢了,没有前途。 他说就没有可以赚快钱的吗?我说有浮动三件套:基金、股票、期货。他说风险太大了,不敢碰。 他说有没有可以躺平就能赚钱的?我说也有躺平三件套:收租、啃老、富二代。他说家庭条件又不是不清楚,哪里有资本? 我说那就只有走捷径三件套:找富婆、认干妈、吃软饭。他说干不了,性格内向,吃不了那碗饭,还跟我生气。 他说没有骨气的事不干,还不死心的问我,难道就没有一夜暴富的吗? 我说有做梦三件套:体彩、福彩、刮刮乐。几十条路摆在你面前,总有一条适合你,他转身就走了,似乎下定了某种决心,再也没有回头看我一眼。社会的毒打会比来自校园的毒打更加刻骨铭心。
我小侄子高考落榜了,考了236分,问我有没有赚钱好的路子? 我说那肯定有,体面三件套。考研、考公、考编,他摇了摇头说学历太低了,考不了。 我说那就牛马三件套。进厂拉货、做销售,他说想自由点的。 我说那就铁人三项。快递、外卖、网约车,他说有没有安稳点的。
GLM-5.3 vs Fable 5 breakdown of DeepSWE performance. Coding quality parity with 1/4th the cost is a sweet deal! > GLM-5.3 at 69.0% vs. 69.7% > $3.99/task vs $21/task > 80k tokens vs 114k tokens > 124 turns vs 85
最近爆火的OpenRouter 上冒出一个匿名模型 Ox Alpha(GLM),其实管道不会撒谎,会撒谎的是人 Crawl4AI的作者unclecode这句话,道破了怎么验一个API背后到底是哪家的模型 厂商在别名后面偷偷换模型,不会告诉你。他两天前就发现DeepSeek的deepseek-chat这个老名字下,悄悄换成了另一个模型 他的思路是不听模型自己说,只看基础设施的指纹,九个探针: 01|喂英文全字母句、中文段落、代码、emoji,只数prompt_tokens。分词器是每家自己训的,中文分词差异最大,藏不住 02|temperature设成2.0,看报错原文逐字。那句话是这家工程师亲手写的 03|GLM的1301错误码、finish_reason的词表、上下文天花板是1048576还是262144,每一项都能定位到具体厂商和变体 前几天OpenRouter上那个匿名模型Ox Alpha,全网都在猜。modelprint一测,分词器四项归一化计数全中,直接指向GLM家,其它厂商最好也只中一半 纯前端无后端,API key只留在你自己的浏览器标签页,不经过任何服务器
吴思先生的《血酬定律》
15分钟一局的王者荣耀,成了数学博士生的避难所。 mp.weixin.qq.com/s/kUWXxWqfryq2…
Read full blog: lmsys.org/blog/2026-08-1…
别人给她台阶是让她下的,不是让她上的。 所以有些人被讨厌是有原因的。
Thoughts About Scaling Law Scaling, but not only of parameters. Every model release now ends with the same question: how many parameters? It isn't a question that can be answered on its own. Parameter count is only meaningful alongside three others — how much data you have, where you intend to spend your compute, and who will run the model, under what conditions. The field learned this the hard way. Kaplan et al. (2020) fit an exponent that told everyone to grow parameters faster than data — roughly 2.7:1 — and the industry complied: GPT-3, Gopher, MT-NLG. Hoffmann et al. (2022) redid the experiment across four hundred models and found the compute-optimal split is closer to 20 tokens per parameter, and that with sufficient compute the two should grow at the same rate rather than drifting apart. The error in the earlier fit compounded with every order of magnitude of compute, which is why the largest models of that generation were the most misallocated. The trillion-parameter round was, in retrospect, a detour the whole field took together and then reversed. Chinchilla wasn't the end either. It optimized training compute for models that would be trained once and evaluated. Today a model is called billions of times a day and inference dominates lifetime cost. Put inference into the objective and the optimum moves toward smaller models trained far longer — deliberate over-training, which is what Llama-2-7B and Gemma-2-9B were doing at roughly 290 and 889 tokens per parameter. Sparsity moved the target again. In a MoE model two quantities have to be kept apart: total parameters govern roughly how much the model can hold — knowledge, facts, the long tail — while activated parameters and effective depth govern roughly how far it can think, how many steps of a causal chain it can carry before it comes apart. A dense 20:1 ratio does not transfer. And the ratio isn't a single number at all: Roberts et al. (2025) find the optimal tokens-per-parameter is task-dependent, with memorization favoring more parameters and reasoning favoring more data. Follow-up work on MoE observes that at fixed TPP, pushing total parameters higher actually degrades reasoning, while activating more experts reliably helps it. This matters for what we are building toward. Finding a vulnerability is not a retrieval problem. It doesn't come from having memorized more CVEs; it comes from carrying a twenty-step chain of inference to the end without losing the thread. That capability does not live in total parameter count. Which brings us to this release. Total parameters appear to matter up to a threshold — enough to hold the world — after which additional capability comes from scaling elsewhere: effective depth per forward pass, and above all post-training. GLM-5.3 is our controlled experiment on that claim. Same base, same architecture, same total and activated parameters as GLM-5.2. One month of scaling long-horizon environments and RL. The gains are not marginal. Well, scaling has more than one dial. We turned the post-training one this time because it had the most slack left in it — not because the others are finished. Base model size, pretraining data, compute spent per forward pass: all of them are still on the table, and we will come back to each. What this experiment taught us is that the dials do not have to be turned together, and that the one worth turning next is rarely the one that was worth turning last. We are not done scaling. Next time, maybe mid-training, pre-training, and even more.
回顾90年代中国改革与“江朱体制” | China's 1990s reforms and the legacy of the 'Jiang-Z... youtu.be/maM7D1BffT0?si… via @YouTube
Ten advances in mathematics and theoretical computer science | OpenAI openai.com/index/ten-adva…
Mcore Bridge:迈向Megatron训练"零门槛"时代_魔搭ModelScope社区-ModelScope魔搭社区 modelscope.csdn.net/691185110e4c46…
何清漣專欄:潘石屹追稅風波——資本跨境套利時代結束 - 鏡報 mirrordaily.news/story/79231
rmrf100 @rmrf100
270 Followers 434 Following
磊哥 @LeiZhaoBuaa
0 Followers 216 Following
🌍 @8l7AZ0Cy4N26949
0 Followers 11 Following
Yasaman Haghighi @yasi_haghighi
2 Followers 25 Following
Andrew Ng CHAT♐♐ @AndrewNgCHAT
21 Followers 373 Following 12following 2.5Mfollowers Co- Founder of Coursera'; Stanford CS adjunct faculty. Former bead of Baidu AI Group/Google Brain. #ai
Tim Karl @karl_tim30191
0 Followers 770 Following
JSmith @jsmith0x7bb
0 Followers 2K Following
Theelyez @Theelyez597b
74 Followers 2K Following
Prof Celso Fontes @profcelsofontes
151 Followers 139 Following
Binin @Binbinin
8 Followers 3K Following
Mpuseasm @MpuseasmoPVr6e
10 Followers 157 Following
HandsomestDayi @HandsomestD
3 Followers 74 Following
Bart @Bart161649
4 Followers 144 Following
xland2023 @xland202352226
58 Followers 3K Following
Meta @MetaMeowMeow
77 Followers 4K Following
Roger @xRog3r
175 Followers 987 Following 🤖 AI 创业 | AI infra | Build in Public 🚀 关注技术,产品和投资 🔗 公众号: Rog3r
zengping @sethbrin
0 Followers 30 Following
Mary @mary_newberry87
170 Followers 3K Following
Jaime Grace @JaimeGrace99076
4 Followers 129 Following Software Engineer | Passionate Gamer | Unity Developer | Night owl 💻🎮
Teaunes @teaunes22388
30 Followers 637 Following
Victor He @ShishengHe
525 Followers 2K Following 建筑与室内设计师|终身学习者|创业探索者 聚焦AI在家居、设计、建造等垂直领域的应用 记录技术演进与行业重构的真实进程 ⚡️AI Living Lab|设计思维 × 算法革命 × 创业洞察
Bondowe Kadiata @bondkad
871 Followers 6K Following un congolais de Kinshasa, Tolingi kimia na Congo. Que les médiocres dégagent. le peuple azalaka kaka zoba.
Ren @rren_nguyen
3 Followers 90 Following Avive Citizen I am an Avive Citizen https://t.co/XNEg0whCai
Sadat Behrami @BehramiSadat
3 Followers 37 Following
Patrick Copeland @copelandpatrick
1K Followers 4K Following Aspiring runway model & astrophotographer.
MarQuis Trill | Youtu... @marquistrillx
1.9M Followers 3.6M Following Bitcoin Class of 2017 | Crypto Trader of The Year 2018 | Binance & Forbes Top 50 Influencer | I manage creators & businesses with content, growth and strategy
CryptoPuppies @CryptoPuppies_
2K Followers 2K Following Collectable. Breedable. Adorable. 🐶 Are Man's Best Friend on the Blockchain. Join our telegram:https://t.co/1ftNDfJLTc #CryptoPuppies
ip2country @itocountry
101 Followers 1K Following
Smart App Marketer @smartappmktr
43 Followers 228 Following Bringing you the latest in mobile app marketing. Follow us out to learn about app marketing & keep up with the latest news.
Grape_424 @Grape_424
139 Followers 1K Following
无 @vpntech
827 Followers 432 Following
Wanli X @Wanli_Xiong
34K Followers 266 Following Graduated from Harvard, Engineer, urban planner, entrepreneur and investor, 工程师、城市规划师、企业家投资人,母校:哈佛、华中科技大学, Life is X, 找一叶没有舵手的船。“德全开万世 鸿明光一朝”之楚王熊家“万”字辈万里,第169代。
CABA下小雪 @2091996981snow
7K Followers 162 Following 旅行者|数字游民|目前定居阿根廷🇦🇷| 大学退学高中学历|知行合一爱国华侨| 去过15个国家|国际低成本旅居专家| 不键政|不可知论|渡自己不渡众生| 备用号@478182920snow 致力于研究整理各国移民法律政策 会员群和移民规划服务请电邮联系 [email protected]
Sand.ai @SandAI_HQ
4K Followers 26 Following Our mission is to advance AI to benefit everyone. @GagaAI_official & @magiai_hq are out now.
Tianyi Cui @tianyi
38K Followers 463 Following Member of Technical Staff @ DeepSeek, Harness Team I'd love to connect with members of international frontier LLM labs! DM is open. Opinions are my own.
Zartbot @zartbotF
7K Followers 122 Following "Chase after the truth like all hell and you'll free yourself, although you never touch its coat tails."
Garry Tan @garrytan
1.1M Followers 6K Following President & CEO @ycombinator —Founder @garryslist—Creator of GStack & GBrain—designer/engineer who helps founders—SF Dem accelerating the boom loop
Mario Zechner @badlogicgames
72K Followers 1K Following Armin's handler at https://t.co/B05ybKGkzx. Old man yelling at Claudes. https://t.co/Q1wG57v1yc https://t.co/mnOoWUr0TO https://t.co/8i5vIRE0Wn
Dmitry Rybin @DmitryRybin1
6K Followers 255 Following Cofounder $100M AI Startup in Shenzhen, Algorithm Discovery + Math (we’re hiring) | ML PhD CUHK, BSc. Math HSE | IMC🥇National Math Olympiad🥇
Aleksa Gordić (水�... @gordic_aleksa
31K Followers 227 Following scaling @ Nemotron getting us to singularity with friends | angel computers can be understood: https://t.co/doHE1Quv2L x @GoogleDeepMind @Microsoft
Andrew Curran @AndrewCurran_
87K Followers 19K Following 🏰 - I write about AI, mostly. Expect some strange sights.
Misha Goin @mgoin_
3K Followers 480 Following maintainer @vllm_project, inference perf @RedHat_AI (acq @neuralmagic)
Lukasz Olejnik @lukOlejnik
32K Followers 266 Following Security & Privacy. Data Protection. Research & Development. Engineering. Analyst. Policy. W3C. Consultant. Author. [email protected] Ph.D, LL.M. @warstudies
Banghua Zhu @BanghuaZ
9K Followers 2K Following Cofounder & CTO @radixark. Assistant Professor @UW. Prior @Nvidia @Berkeley_EECS
KVCache.AI @KVCache_AI
1K Followers 111 Following Hi, this is https://t.co/EO7MXLjjSU official account. We build systems for efficient LLM serving, including KTransformers, Mooncake and AgentENV.
Hunter Bown @goodhunt
11K Followers 6K Following Creator & maintainer of Codewhale. Issues & PRs always welcome. https://t.co/bT1r4IQHqc
Mao Keji | मुख�... @kejimao
15K Followers 396 Following Analyst at the International Cooperation Center, Visiting Fellow at Harvard Yenching Institute (Views are my own)
Deli Chen @victor207755822
32K Followers 181 Following Deep Learning Researcher @deepseek_ai | #AGIforEveryone Prev. BS and MS @PKU1898 | https://t.co/nu6M0PNxoM | All opinions are my own. | INTP-T | 人心惟危,道心惟微
Xiaokang Chen @PKUCXK
6K Followers 63 Following Researcher in Multimodal Team @deepseek_ai Previously Bachelor & Ph.D at Peking University @PKU1898 Opinions are my own.
Dwarkesh Patel @dwarkesh_sp
269K Followers 1K Following Host of @dwarkeshpodcast https://t.co/3SXlu7fy6N https://t.co/4DPAxODFYi https://t.co/hQfIWdM1Un
CrazyBoyM @baicai003
3K Followers 345 Following Fcking. Accelerating. Hacking the World. 🌍 Founder of ShareAI Lab Building something for agent
(((ل()(ل() 'yoav)))... @yoavgo
87K Followers 2K Following
Chen Sun 🤖 @ChenSun92
3K Followers 445 Following Research Scientist @GoogleDeepMind RSI, memory, automated science ex-IMO (🇨🇦) ex-neuroscientist Views are my own
邓聿文 @dyw1968316
123K Followers 220 Following Analyst on CCP elite politics, foreign policy & Taiwan issues |精英政治、外交与两岸|FA published|FP & ThinkChina contributor|Former political advisor, Eurasia Group 油管频道
陳軍 @chenjunnyc
33K Followers 5K Following
Kimi.ai @Kimi_Moonshot
359K Followers 137 Following Built by Moonshot AI to empower everyone to be superhuman. PR: [email protected] DC: https://t.co/wBsBTn6fmE
傅盛 @FuSheng_0306
56K Followers 312 Following 猎豹移动董事长兼CEO,猎户星空董事长@orionstar2016 | 生日3月6日,打造了早期的360安全卫士|TikTok最早期的投资人|是上市公司CEO,也是陪女儿打游戏的老爸玩家| 是连续创业者,也是从220斤爆改到150斤的运动达人
zhyncs @zhyncs42
4K Followers 1K Following LightSeek Mafia 🌁 OPINIONS ARE MY OWN, Senior Director @togethercompute, Governing Board @lightseekorg, Homepage https://t.co/saCowtppUm
yan5xu @yan5xu
17K Followers 480 Following 🤖 AI 野生研究员 | ex @ManusAI_HQ & @hey_im_monica 推特内容仅代表个人观点,和公司无关
Chi Jin @chijinML
9K Followers 524 Following Researcher @OpenAI | Associate Prof @Princeton AI Reasoning · Reinforcement Learning · Game Theory · ML Foundations
Ant Ling @AntLingAGI
13K Followers 5 Following MoE model series with foundation (Ling), reasoning (Ring) and any-to-any (Ming) from Ant Group’s AGI initiative, @TheInclusionAI. https://t.co/6LEkFlo2cq
vLLM @vllm_project
48K Followers 36 Following A high-throughput and memory-efficient inference and serving engine for LLMs. Join https://t.co/lxJ0SfX5pJ to discuss together with the community!
Zhiqing Sun @EdwardSun0909
20K Followers 1K Following Lead agent research @Meta MSL TBD Lab. previously posttraining/agent research @OpenAI. CS PhD @LTIatCMU
Ruiqi Gao @RuiqiGao
15K Followers 884 Following @AnthropicAI | Prev. @google brain/deepmind | Mom of Mochi.
Papers of the day @ArxivToday
1K Followers 1 Following Best papers from @arxiv, maintained by @ennucore and LLMs
yanguoliusheng @szygls
75K Followers 20K Following Current events and gossip; love for the country; friendly discussion; no profanity.
Shaun Rein @shaunrein
102K Followers 1 Following Founder of The China Market Research Group (CMR). Author of 5 books. The Split: Finding the Opportunities in China's Economy in the New World Order @harvard
Nando de Freitas @NandoDF
110K Followers 909 Following I seek to understand intelligence & agency and build AI aligned with compassion, freedom & universal human empowerment through progress in science & engineering
Alec Helbling @alec_helbling
11K Followers 2K Following Interpretability, Multimodality, Diffusion. PhDing @GeorgiaTech. NSF Fellow. Prev intern @Apple, @Adobe, @NASAJPL.
Zheng Yuan @GanjinZero
2K Followers 841 Following Seed-Prover, Lean-Workbook, RRHF, RFT and MATH-Qwen. Prev @Alibaba_Qwen, Phd at @Tsinghua_Uni







































