Topic

#open-weight-models

35 articles tagged open-weight-models. Browse the full set below, or see all topics.

Tagged "open-weight-models"

Cross-cutting reads on this topic

35 articles
Mistral is the only frontier-class vendor whose EU answer is structural: Paris HQ, announced EU inference capacity, open weights. When the premium pays.
#mistral-ai#sovereign-ai+5 more
2026-08-03
Read Article
Qwen3.8-Max ships as a 2.4T-parameter MoE with 95B active and a 1M context. What the vendor benchmark table shows, and what is still only promised.
#qwen#alibaba+5 more
2026-08-03
Read Article
An industry letter on open weights grew from 25 to 50 signatories in two days. Behind it sit sanctions threats and quieter federal procurement levers.
#open-weight-models#ai-policy+5 more
2026-07-26
Read Article
Ant Group's Ling-3.0-flash is a 124B MoE with 5.1B params. Ant says it matches its 1T flagship; free on OpenRouter to Aug 3, but weights aren't posted yet.
#ling-3-0-flash#ant-group+5 more
2026-07-24
Read Article
OpenRouter added 10 frontier and open-weight models in 22 days, prices spanning ~83x on input and ~250x on output. The July 2026 releases worth routing to.
#openrouter#llm models+4 more
2026-07-23
Read Article
Meituan, the food-delivery super-app, now ships LongCat 2.0 — a 1.6T open-weight (MIT) model trained on Chinese chips. Ant Group and Tencent are doing the same.
#meituan-longcat-2-0#open-weight-models+5 more
2026-07-22
Read Article
Hugging Face says an autonomous AI agent — not a human — ran an end-to-end intrusion of its infrastructure, stealing internal datasets and credentials.
#hugging-face#ai-security+5 more
2026-07-20
Read Article
Alibaba previewed Qwen3.8-Max at WAIC Shanghai, claiming 2.4 trillion parameters and second only to Fable 5 — yet shipped zero benchmarks to back it.
#qwen#alibaba+5 more
2026-07-20
Read Article
Kimi K3's open weights are promised by July 27, not shipped — use the 10-day window to prep hosting for a 2.8T model and check the license before you commit.
#kimi k3#open-weight models+5 more
2026-07-17
Read Article
Kimi K3's open-weight bet meets Anthropic's closed Fable 5. Vendor benchmarks split 6-8, K3 lists at $3/$15 vs $10/$50, and weights are promised July 27.
#kimi k3#claude fable 5+5 more
2026-07-17
Read Article
Five open-weight moves cluster in one July window — K3, Inkling, M3 Pro, a Mistral MoE teaser and a scheduled DeepSeek V4 — narrowing the gap to a generation.
#open-weight models#kimi k3+5 more
2026-07-17
Read Article
Kimi K3 brings 2.8T parameters, 1M context, and near-frontier vendor benchmarks, with open weights due July 27, 2026. What the release means for AI buyers.
#kimi k3#moonshot ai+5 more
2026-07-17
Read Article
Mira Murati's Thinking Machines shipped Inkling, a 975B-parameter Apache 2.0 model built to be fine-tuned, not to top benchmarks. The customize-don't-rent bet.
#inkling#thinking-machines+6 more
2026-07-16
Read Article
Tencent shipped Hunyuan Hy3, a 295B open-weight MoE, under Apache 2.0 on July 6, 2026. It trails GLM-5.2 on coding but wins on agentic tasks at half the memory.
#hunyuan-hy3#tencent+5 more
2026-07-06
Read Article
Pair Fable 5 as the orchestrator brain with GLM-5.2 as open-weight execution muscle. A concrete routing recipe, the cost math, and the single-model risk case.
#fable-5#glm-5-2+4 more
2026-07-03
Read Article
GLM-5.2 is MIT-licensed and downloadable, but full weights run ~1.5TB and need datacenter GPUs. The hardware reality, quant ladder, and routes that work.
#glm-5-2#self-hosting-llms+4 more
2026-07-03
Read Article
Match 2026's best open-weight coding models to your hardware: Qwen3-Coder-Next, Devstral 2, GLM-5.2 and DeepSeek V4 by VRAM, SWE-bench score and real speed.
#Open-Weight Models#Coding Models+6 more
2026-06-29
Read Article
US frontier-model gating is pushing global builders toward open-weight models China already dominates. Does Washington's lockdown hand Beijing the edge?
#China AI#Open-Source AI+6 more
2026-06-27
Read Article
On June 26, 2026 the US gated two frontier models the same day: Anthropic's Mythos 5 and OpenAI's GPT-5.6 Sol. Inside the new managed-release era.
#Government-Gated AI#AI Regulation+6 more
2026-06-27
Read Article
The Fable 5 export shutdown showed single-vendor AI can halt your business overnight. A four-step second-source playbook with open-weight failover backups.
#AI vendor resilience#open-weight models+5 more
2026-06-21
Read Article
DiffusionGemma is Google's first open-weight text diffusion LLM: a 26B MoE under Apache 2.0 hitting 1,100+ tokens/sec on one H100. Where it wins and loses.
#diffusiongemma#google-deepmind+5 more
2026-06-13
Read Article
OpenRouter added five models in ten days, from Opus 4.8 to MiniMax M3 at $0.30/M input. June 2026 pricing, context windows, and usage rankings.
#openrouter#minimax-m3+6 more
2026-06-04
Read Article
Open-weight models now run 10-12x cheaper than frontier SaaS. A 2026 framework on the TCO crossover, vendor lock-in, and data sovereignty for agencies.
#build-vs-buy#ai-strategy+6 more
2026-06-04
Read Article
MiniMax M3 lands at 5-17x lower cost, but Opus 4.8 leads SWE-bench Pro and GPT-5.5 wins Terminal-Bench. A full three-way agentic coding routing matrix.
#minimax-m3#claude-opus-4-8+6 more
2026-06-03
Read Article
DeepSeek abandons its no-outside-capital stance in a ~$7.4B maiden round led by Tencent and CATL, valuing it near $59B and reshaping open-weight economics.
#deepseek#ai-funding+6 more
2026-06-03
Read Article
Cosmos 3 is the first fully open physical-AI omnimodel: one model reasons, simulates, and predicts robot actions. Inside the two-tower design and how to run it.
#nvidia-cosmos-3#physical-ai+6 more
2026-06-01
Read Article
MiniMax M3 fuses frontier coding, a 1M-token context window, and native multimodality. Inside its Sparse Attention design, vendor benchmarks, and pricing.
#minimax-m3#open-weight-models+6 more
2026-05-31
Read Article
StepFun's Apache-2.0 Step 3.7 Flash pairs a 196B MoE backbone with a 1.8B vision encoder, activating ~11B params per token. The cost case for agentic teams.
#stepfun-step-3-7-flash#mixture-of-experts+6 more
2026-05-30
Read Article
When self-hosting open-weight models beats API calls: a cost-crossover model, GPU sizing tables, and a deployment matrix for vLLM, SGLang, and Ollama.
#self-hosting-llm#open-weight-models+6 more
2026-05-27
Read Article
Migrate DeepSeek V3.2 to V4 across open-weight stacks — three reasoning modes, tokenizer change, HCA/CSA attention deltas, KV-cache reduction.
#deepseek-v4#migration-playbook+7 more
2026-05-05
Read Article
GPU spend, ops headcount, latency, and break-even volume for hosting Llama, Qwen, DeepSeek, and Mistral yourself vs API. With per-token cost curves at 4 scales.
#self-hosting-llm#ai-tco+8 more
2026-04-24
Read Article
Cross-model quality regression, throughput lift, and VRAM savings at GPTQ-4, AWQ-4, INT8, and FP8 — benchmark data across 6 open-weight models.
#quantization#gptq+8 more
2026-04-24
Read Article
Seven serverless inference providers compared on price, latency, model availability, and throughput. 60+ data points across 12 popular models.
#ai-inference-providers#together-ai+8 more
2026-04-24
Read Article
Side-by-side input, output, cached, and batch pricing for 30 frontier and open-weight models across 12 providers. Updated April 2026 with 200+ price points.
#ai-model-pricing#llm-pricing+8 more
2026-04-23
Read Article
Q2 2026 gap analysis between open-weight and closed-source frontier models — capability parity, cost economics, and the agency deployment decision tree.
#open-weight-models#closed-source-models+4 more
2026-04-12
Read Article