Topic

#open-weight-models

26 articles tagged open-weight-models. Browse the full set below, or see all topics.

Tagged "open-weight-models"

Cross-cutting reads on this topic

26 articles
Kimi K3's open weights are promised by July 27, not shipped — use the 10-day window to prep hosting for a 2.8T model and check the license before you commit.
#kimi k3#open-weight models+5 more
2026-07-17
Read Article
Kimi K3's open-weight bet meets Anthropic's closed Fable 5. Vendor benchmarks split 6-8, K3 lists at $3/$15 vs $10/$50, and weights are promised July 27.
#kimi k3#claude fable 5+5 more
2026-07-17
Read Article
Five open-weight moves cluster in one July window — K3, Inkling, M3 Pro, a Mistral MoE teaser and a scheduled DeepSeek V4 — narrowing the gap to a generation.
#open-weight models#kimi k3+5 more
2026-07-17
Read Article
Kimi K3 brings 2.8T parameters, 1M context, and near-frontier vendor benchmarks, with open weights due July 27, 2026. What the release means for AI buyers.
#kimi k3#moonshot ai+5 more
2026-07-17
Read Article
Tencent shipped Hunyuan Hy3, a 295B open-weight MoE, under Apache 2.0 on July 6, 2026. It trails GLM-5.2 on coding but wins on agentic tasks at half the memory.
#hunyuan-hy3#tencent+5 more
2026-07-06
Read Article
Pair Fable 5 as the orchestrator brain with GLM-5.2 as open-weight execution muscle. A concrete routing recipe, the cost math, and the single-model risk case.
#fable-5#glm-5-2+4 more
2026-07-03
Read Article
GLM-5.2 is MIT-licensed and downloadable, but full weights run ~1.5TB and need datacenter GPUs. The hardware reality, quant ladder, and routes that work.
#glm-5-2#self-hosting-llms+4 more
2026-07-03
Read Article
Match 2026's best open-weight coding models to your hardware: Qwen3-Coder-Next, Devstral 2, GLM-5.2 and DeepSeek V4 by VRAM, SWE-bench score and real speed.
#Open-Weight Models#Coding Models+6 more
2026-06-29
Read Article
US frontier-model gating is pushing global builders toward open-weight models China already dominates. Does Washington's lockdown hand Beijing the edge?
#China AI#Open-Source AI+6 more
2026-06-27
Read Article
On June 26, 2026 the US gated two frontier models the same day: Anthropic's Mythos 5 and OpenAI's GPT-5.6 Sol. Inside the new managed-release era.
#Government-Gated AI#AI Regulation+6 more
2026-06-27
Read Article
The Fable 5 export shutdown showed single-vendor AI can halt your business overnight. A four-step second-source playbook with open-weight failover backups.
#AI vendor resilience#open-weight models+5 more
2026-06-21
Read Article
DiffusionGemma is Google's first open-weight text diffusion LLM: a 26B MoE under Apache 2.0 hitting 1,100+ tokens/sec on one H100. Where it wins and loses.
#diffusiongemma#google-deepmind+5 more
2026-06-13
Read Article
OpenRouter added five models in ten days, from Opus 4.8 to MiniMax M3 at $0.30/M input. June 2026 pricing, context windows, and usage rankings.
#openrouter#minimax-m3+6 more
2026-06-04
Read Article
Open-weight models now run 10-12x cheaper than frontier SaaS. A 2026 framework on the TCO crossover, vendor lock-in, and data sovereignty for agencies.
#build-vs-buy#ai-strategy+6 more
2026-06-04
Read Article
MiniMax M3 lands at 5-17x lower cost, but Opus 4.8 leads SWE-bench Pro and GPT-5.5 wins Terminal-Bench. A full three-way agentic coding routing matrix.
#minimax-m3#claude-opus-4-8+6 more
2026-06-03
Read Article
DeepSeek abandons its no-outside-capital stance in a ~$7.4B maiden round led by Tencent and CATL, valuing it near $59B and reshaping open-weight economics.
#deepseek#ai-funding+6 more
2026-06-03
Read Article
Cosmos 3 is the first fully open physical-AI omnimodel: one model reasons, simulates, and predicts robot actions. Inside the two-tower design and how to run it.
#nvidia-cosmos-3#physical-ai+6 more
2026-06-01
Read Article
MiniMax M3 fuses frontier coding, a 1M-token context window, and native multimodality. Inside its Sparse Attention design, vendor benchmarks, and pricing.
#minimax-m3#open-weight-models+6 more
2026-05-31
Read Article
StepFun's Apache-2.0 Step 3.7 Flash pairs a 196B MoE backbone with a 1.8B vision encoder, activating ~11B params per token. The cost case for agentic teams.
#stepfun-step-3-7-flash#mixture-of-experts+6 more
2026-05-30
Read Article
When self-hosting open-weight models beats API calls: a cost-crossover model, GPU sizing tables, and a deployment matrix for vLLM, SGLang, and Ollama.
#self-hosting-llm#open-weight-models+6 more
2026-05-27
Read Article
Migrate DeepSeek V3.2 to V4 across open-weight stacks — three reasoning modes, tokenizer change, HCA/CSA attention deltas, KV-cache reduction.
#deepseek-v4#migration-playbook+7 more
2026-05-05
Read Article
GPU spend, ops headcount, latency, and break-even volume for hosting Llama, Qwen, DeepSeek, and Mistral yourself vs API. With per-token cost curves at 4 scales.
#self-hosting-llm#ai-tco+8 more
2026-04-24
Read Article
Cross-model quality regression, throughput lift, and VRAM savings at GPTQ-4, AWQ-4, INT8, and FP8 — benchmark data across 6 open-weight models.
#quantization#gptq+8 more
2026-04-24
Read Article
Seven serverless inference providers compared on price, latency, model availability, and throughput. 60+ data points across 12 popular models.
#ai-inference-providers#together-ai+8 more
2026-04-24
Read Article
Side-by-side input, output, cached, and batch pricing for 30 frontier and open-weight models across 12 providers. Updated April 2026 with 200+ price points.
#ai-model-pricing#llm-pricing+8 more
2026-04-23
Read Article
Q2 2026 gap analysis between open-weight and closed-source frontier models — capability parity, cost economics, and the agency deployment decision tree.
#open-weight-models#closed-source-models+4 more
2026-04-12
Read Article