Topic

#mixture-of-experts

13 articles tagged mixture-of-experts. Browse the full set below, or see all topics.

Tagged "mixture-of-experts"

Cross-cutting reads on this topic

13 articles
Qwen3.8-Max ships as a 2.4T-parameter MoE with 95B active and a 1M context. What the vendor benchmark table shows, and what is still only promised.
#qwen#alibaba+5 more
2026-08-03
Read Article
Thinking Machines' 276B/12B Inkling-Small leads its 975B parent on the vendor's own coding and reasoning table, then loses badly on factuality.
#open-weights#thinking-machines+5 more
2026-08-01
Read Article
Ant Group's Ling-3.0-flash is a 124B MoE with 5.1B params. Ant says it matches its 1T flagship; free on OpenRouter to Aug 3, but weights aren't posted yet.
#ling-3-0-flash#ant-group+5 more
2026-07-24
Read Article
Meituan, the food-delivery super-app, now ships LongCat 2.0 — a 1.6T open-weight (MIT) model trained on Chinese chips. Ant Group and Tencent are doing the same.
#meituan-longcat-2-0#open-weight-models+5 more
2026-07-22
Read Article
poolside's 118B open-weight Laguna S 2.1 runs on one DGX Spark and, on its own harness, tops rivals 10x its size on Terminal-Bench — trained in under 9 weeks.
#poolside-laguna-s-2-1#open-weight-coding-model+5 more
2026-07-22
Read Article
Mira Murati's Thinking Machines shipped Inkling, a 975B-parameter Apache 2.0 model built to be fine-tuned, not to top benchmarks. The customize-don't-rent bet.
#inkling#thinking-machines+6 more
2026-07-16
Read Article
Cohere's first open-source model is a 30B MoE that runs on a single H100, scores 33.4 on the Coding Index, and ships under Apache 2.0. Full breakdown.
#cohere#north-mini-code+5 more
2026-06-13
Read Article
NVIDIA shipped Nemotron 3 Ultra, a 550B open MoE reasoning model with weights, data and recipes under a permissive license. It runs fast but trails Kimi K2.6.
#nvidia#nemotron+6 more
2026-06-05
Read Article
StepFun's Apache-2.0 Step 3.7 Flash pairs a 196B MoE backbone with a 1.8B vision encoder, activating ~11B params per token. The cost case for agentic teams.
#stepfun-step-3-7-flash#mixture-of-experts+6 more
2026-05-30
Read Article
DeepSeek-V4 ships April 24, 2026 as open-weight MoE: Pro (1.6T/49B active) and Flash (284B/13B), 1M context, 27% FLOPs and 10% KV cache vs V3.2.
#deepseek-v4#deepseek-v4-pro+6 more
2026-04-24
Read Article
MoE choices powering 2026 frontier models compared — total vs active params, routing strategies, sparsity ratios, and the downstream cost implications.
#mixture-of-experts#moe-architecture+8 more
2026-04-24
Read Article
Mistral Small 4 packs 119B parameters (6B active via MoE) with 256K context under Apache 2.0. Unified model for reasoning, vision, and coding.
#mistral-small-4#mixture-of-experts+5 more
2026-03-16
Read Article
Qwen 3.5 medium series: Flash, 35B-A3B, 122B-A10B, and 27B. Benchmarks vs GPT-5 mini and Claude Sonnet 4.5, pricing from $0.10/M tokens.
#qwen-3-5#alibaba-ai+6 more
2026-02-25
Read Article