AI DevelopmentReference7 min readPublished September 21, 2026

16 rows · 8 licences · 1 fully open training release · the licence file, not the repo tag

What "Open Source" Actually Includes for an AI Model

16 models in the open-model conversation, checked against the licence file, OSI's list and what else was released: weights, data, recipes, evaluation harness.

DA
Digital Applied Team
Research and practical guidance
Editorial dateSeptember 21, 2026
Licence files readSeptember 22, 2026

"Open source" on a model page can mean a downloadable file under any terms, a licence the Open Source Initiative lists, or a release complete enough to rebuild. Those are three different things, and a model can have any one of them without the other two. A team that depends on an open model as its fallback needs to know which of the three it actually has.

This census reads the primary for each: the licence file rather than the repository tag, the Open Source Initiative's own list rather than a vendor's adjective, and the model card and repository for what else was published. Sixteen rows covering nine publishers are in the grid. Where a publisher does not say something, the cell says "not stated", which is not a synonym for no.

Key takeaways
  1. 01
    Fourteen of sixteen have downloadable weights. Seven carry a licence on the OSI list. One publishes enough to rebuild.The two without weights are the newest variants of models that do have them: GLM-5.3-FlashX and Muse Spark 1.3. The one with data, recipes and a pinned evaluation harness is NVIDIA's Nemotron, and its licence is not on the OSI list.
  2. 02
    Five bespoke licences restrict on revenue or user count, not on subject matter.Qwen3.8-Max at US$50M for model-as-a-service businesses, Qwen Community at any revenue, Kimi K3 at US$20M, LFM Open at US$10M for any commercial use, Llama 4 at 700M monthly active users. They are commercial terms wearing a licence's name.
  3. 03
    Permissive is not the same as OSI-listed, and a family name tells you nothing.OpenMDW-1.1 grants rights without restriction and is absent from the OSI list. Qwen ships one model under Apache-2.0 and two under bespoke terms, and the tighter of the two is on the smaller model.
  4. 04
    A free route proves a hosting arrangement, not a licence.Of 21 free routes on one catalog, three run models with no public weights at all, three carry expiry dates within the month, and the two unmodified Apache-2.0 gpt-oss models have no free route.

01DefinitionsThree words, three tests

Each term has a test a reader can run, and each test looks at a different artefact. Run all three before calling anything open.

Open weights
Can you download the parameters?
14of 16

The artefact, under whatever terms the publisher chose. Says nothing about the licence. Fails for GLM-5.3-FlashX and Muse Spark 1.3, both reachable only through a hosted API.

Test: a public repository
Open source
Is the licence in the LICENSE file on the OSI list?
7of 16

Use, modify, redistribute and sell without a threshold, a permission step or a field-of-use carve-out. Fails for five bespoke licences and for OpenMDW, the licence on the most complete release in the grid.

Test: the file, then the OSI list
Open training
Are data, recipes, environments and harness published?
1of 16

Enough to rebuild or re-verify. Nemotron publishes datasets including reinforcement-learning sets, training recipes and evaluation recipes with pinned containers. DeepSeek-V4.1-Flash publishes benchmark-reproduction steps; every other row leaves at least one of the four not stated as a release.

Test: the repository and the card

The Open Source Initiative check used here is the defensible one: does the licence's identifier appear in OSI's published list, and does its page on that site resolve. Apache-2.0 and MIT pass both. OpenMDW, the two Qwen licences, the Kimi K3 licence, the LFM Open License and the Llama 4 licence fail both. The list does not distinguish a licence that was never submitted from one under review, so "not listed" is all this page claims.

02The censusThe licence grid

Every licence cell below comes from the licence file in the weights repository or the publisher's own licence page, read in full. Two exceptions are marked: Qwen3.8-27B's Apache-2.0 comes from card metadata because the file fetch returned a redirect, and Inkling's Apache-2.0 is declared in card metadata with no licence file in the repository. Two rows have no licence file at all, and that is recorded rather than filled.

Sources: each repository's LICENSE file or the publisher's licence page, and OSI's licence list, read on the as-of date in the methodology. Thresholds paraphrase the licence clauses; the files are linked in the text.
ModelLicenceRestricts use?OSI-listed?
GLM-5.3-Flash (Z.ai)MITNoYes
GLM-5.3-FlashX (Z.ai)No weights publishedn/an/a
DeepSeek-V4.1-FlashMITNoYes
DeepSeek-V4-Pro-0813MITNoYes
Qwen3.8-2.4T-A95B (Max-class)Qwen3.8-Max LicenseYes: separate licence for model-as-a-service businesses over US$50M revenue; attribution over 100M MAU or US$20M monthly revenueNo
Qwen3.8-Flash-NextQwen Community License 1.0Yes: separate licence for any commercial model-as-a-service use, no revenue floor; same attribution clauseNo
Qwen3.8-27BApache-2.0 (card metadata)NoYes
gpt-oss-120b / 20b (OpenAI)Apache-2.0No; a two-sentence usage policy requires compliance with applicable lawYes
Nemotron 3 Ultra 550B (NVIDIA)OpenMDW-1.1No field-of-use limit; rights end if you sue over the model materialsNo
Nemotron 3.5 Lightning 30B (NVIDIA)OpenMDW-1.1SameNo
Muse Glimmer 30B (Meta)Apache-2.0, plus a separate usage policyNot in the licence; the policy's standing is not statedYes, licence only
Muse Spark 1.3 / 1.3-Contributor (Meta)No weights publishedn/an/a
Llama 4 Scout / Maverick (Meta)Llama 4 Community LicenseYes: permission required over 700M monthly active users; "Built with Llama" attribution; derived models must carry "Llama" in the nameNo
Kimi K3 (Moonshot AI)Kimi K3 LicenseYes: separate agreement for model-as-a-service businesses over US$20M revenue; attribution clauseNo
Inkling (Thinking Machines)Apache-2.0 declared; no licence file in the repoNot in the licence; a separate acceptable-use policy is linkedYes, as named
LFM2.5-2.6B (Liquid AI)LFM Open License v1.0Yes: commercial use by an entity with US$10M or more annual revenue is not licensedNo

Four rows deserve a sentence each. Z.ai's GLM-5.3-Flash licence is unmodified MIT, which is as clean as the grid gets; Z.ai's own documentation calls the model "the first open-source frontier model to combine sparse and linear attention", and that is Z.ai's claim, quoted, not ours. The Qwen3.8-Max licence requires a separate licence from Qwen before any commercial use by a model-as-a-service or AI work-assistant business whose group revenue passes US$50 million in any twelve months; the Community licence on Qwen3.8-Flash-Next has the same clause with no revenue floor. NVIDIA's OpenMDW-1.1 grants permission to deal in the model materials without restriction and says it imposes no obligations on outputs, and it is still not on the OSI list. Meta's Llama 4 licence requires a licensee above 700 million monthly active users to request permission that Meta may grant at its sole discretion.

Two rows pair a permissive licence with a separate document. Muse Glimmer ships unmodified Apache-2.0 beside a usage policy with a prohibited-uses list and an under-18 exclusion; Inkling declares Apache-2.0 in card metadata with no licence file in the repository and links an acceptable-use policy. Neither publisher states how the second document relates to the licence, so this page records both files and adjudicates nothing. Muse Spark 1.3, the closed model Muse Glimmer is distilled from, has no licence file because it has no weights; what Meta's own pricing table does state is that the cheaper Contributor tier's traffic is "Used to improve our products" at $0.10 input and $0.20 output per million tokens, while the standard tier at $1.25 and $4.25 is "Not used to improve our products". The discount is paid in data.

Any Commercial Use of the Work or a Derivative Work by a Legal Entity that exceeds the Threshold is not licensed under this Agreement.LFM Open License v1.0, section 5(b), where the threshold is annual revenue of US$10 million or more

03Beyond the licenceWhat else was released

A licence governs the weights. Whether anyone could rebuild or re-verify the model depends on what else the publisher put in the repository. Nine publishers in seven lines, from their own cards and repositories; the Nemotron 3.5 Lightning card and the gpt-oss-120b card are the two ends of the range.

NVIDIA NemotronCard states "open weights, training data, and recipes". 40+ Nemotron dataset repositories including RL sets; evaluation recipes published with pinned containers, prompts and scoring; a technical report; the card also lists which benchmarks are not yet reproducible.
Data, recipes, harness
DeepSeek-V4.1-FlashTechnical report inside the weights repo; an evaluation folder with step-by-step DeepSWE v1.1 reproduction; data described by size and shape only (45T tokens); no training code or environments.
Report, eval steps
Z.ai GLM-5.3-FlashGLM-5 series technical report on arXiv; harnesses named with versions in benchmark footnotes; a 30T-token corpus stated by size; no training code, environments or dataset release.
Report, harness names
OpenAI gpt-ossModel card paper on arXiv; repository holds inference, tooling and format support. The card says nothing about the training corpus: no size, no sources.
Report only
Qwen3.8 familyFlash-Next has a technical report in its GitHub repository; the Max-class card links a blog post. Training data, code and environments not stated on any of the three cards.
Report (one model)
Meta Muse GlimmerTraining data described in one sentence: public data, third-party data and information from Meta's products and services. No size, no sources, no report for the model itself.
One sentence on data
Kimi K3, Inkling, LFM2.5Kimi K3 ships a technical report; Inkling describes its data in prose with no numbers and dates its comparison run; LFM2.5 cites the predecessor line's report. Training code and environments not stated for all three.
Report or prose
The two-way test

Nemotron is the most complete training release in the grid and is not open source by the OSI test. gpt-oss is unambiguously open source and publishes nothing about its data. The licence badge and the completeness claim measure different things. A team that conflates them will get one of the two wrong, and which one depends on what they needed the model for.

04The catalogWhat a free route tells you

A router's free variant of a model looks like an openness signal. On the day of writing OpenRouter carried 21 routes with a free suffix, and reading them against the publishers' repositories gives four flat statements.

A free route does not imply open weights: three of the 21 run models with no public weights anywhere. A free route does not imply a permissive licence: one runs LFM2.5, whose licence withdraws commercial rights above US$10 million in revenue. A permissive licence does not produce a free route: both gpt-oss models are unmodified Apache-2.0 and every gpt-oss route on the catalog is paid. And a free route can expire: three of the 21 carried deprecation dates in September, which OpenRouter documents as the deprecation date for the model endpoint. A free route is a hosting arrangement with an end date. Our price index tracks the paid routes with their check dates.

05The decisionThree questions before a second source

The case for an open model in a production stack is usually resilience: if the hosted vendor changes terms or disappears, the open model is the fallback. Our second-source playbook covers the architecture; these three questions decide whether a given model qualifies, and each maps to a column in the grid.

Which file is the licence, and what does its commercial clause say?
Read the LICENSE file or the vendor's licence page, not the repository tag or the card. Four of sixteen rows carry a use restriction in the licence file that the model page does not mention, and four rows have no licence file in the weights repository at all.
The licence file
Is the newest variant the one with weights?
Two rows fail: GLM-5.3-FlashX has no public repository, and Muse Spark 1.3 has none while the open Muse Glimmer is distilled from it. A third, Qwen3.8-Max, has open base weights, but the shipping product adds vision, a 1M default context and built-in tools the weights do not carry. A fallback one tier behind what you run is not a fallback.
The repository
If the hosted route vanished tomorrow, could you run it and re-verify it?
Weights answer the first half. Published data, recipes, environments and pinned evaluation recipes answer the second. One row in the grid answers both; if yours does not, budget for your own evaluation harness before you need it.
The card and the repo

For the hardware side of the same decision, which open models fit which machines, see our self-hosting guide. If you want the licence review and the fallback design done together, that is part of what we do under AI transformation.

06MethodologyHow this page was built

Methodology

Primary sources only: licence files, model cards, launch posts, technical reports, registry metadata and the Open Source Initiative's own licence list. No press article or aggregator is cited.

What was collected
For 16 rows from nine publishers: whether weights are public and where; the licence name from the licence file; whether it restricts use and how; whether its identifier is on OSI's list; whether training code, post-training environments, an evaluation harness, a data description and a technical report are published; and whether a hosted API is the only route to the newest variant. Plus all 21 free routes on one catalog.
Sources
Hugging Face repositories for Z.ai, DeepSeek, Qwen, OpenAI, NVIDIA, Meta, Moonshot AI, Thinking Machines and Liquid AI, including each LICENSE file and model card; openmdw.ai for OpenMDW-1.1; dev.meta.ai for the Llama 4 licence and the Muse Spark pricing table; research.meta.ai for the Muse Glimmer launch post of August 10, 2026; Z.ai's documentation with a Wayback capture of August 26, 2026; NVIDIA's NeMo Gym and Evaluator repositories; the OSI licence API; OpenRouter's model list and documentation.
As-of date
All licence files, cards, catalog values and the OSI list were read on September 22, 2026. The page is dated September 21 for the week it covers; undated vendor pages are quoted as they stood on the reading date, not asserted to have been identical on September 21.
Units
Revenue and user thresholds are as written in each licence. Parameter counts are the publisher's own. "OSI-listed" means the identifier appears in OSI's published list and its page on the Open Source Initiative's site resolves; OSI's API carries an approval flag that tracks its current review process and was not used.
Exclusions
Any model whose vendor release note is dated after September 21, 2026. Google Gemma 4, IBM Granite, Cohere, Poolside, MiniMax, StepFun, Tencent and others were present in the catalog but not read to grid depth, so they are left out rather than half-filled. Llama 4's data description and report were not checked and are not claimed.
Known limitations
Qwen3.8-27B's licence is from card metadata, not a read file. Inkling's Apache-2.0 is declared with no in-repo file. Whether a separate usage policy binds alongside a permissive licence is not stated by any publisher and is not resolved here. OSI's list does not distinguish never-submitted from under-review.
Refresh
Refreshed in place when a listed model changes licence, when a newer variant of a listed model publishes or withholds weights, or when OSI lists a licence in the grid.

07ConclusionOpen weights, open source and open training are independent, and the grid shows it both ways

What to do with this

Read the licence file, check the newest variant, and decide what you would need if the hosted route went away

Fourteen of these sixteen rows can be downloaded. Seven name a licence the Open Source Initiative lists, though two of those seven pair it with a separate usage policy or declare it with no file in the repository. One can be rebuilt and re-verified from what its publisher released. Before treating any of them as a second source, find the licence file and read its commercial clause, confirm the weights you can get are the model you would actually run, and write down what you would need to evaluate it yourself. The grid gives the answers for these sixteen; the three tests work for the next sixteen.

Digital Applied

A fallback model you have actually checked.

We review model licences, weights availability and reproducibility for teams building a second source into their AI stack, and we keep this grid current as publishers change terms.

Licence reviewSecond-source designEvaluation harness
Your next project

An open model you can depend on

  • Which licence you are actually under
  • Whether the weights match the product
  • What you need to re-verify it yourself
Questions and answers

Applying this post

Not on that evidence. The page proves weights are downloadable. Whether the model is open source depends on the LICENSE file in the repository or the publisher's licence page, and whether that licence is on the Open Source Initiative's list. Five of the sixteen rows here carry a use restriction in the licence file, and four have no licence file in the weights repository at all.
Digital Applied newsletter

Deep dives on AI, marketing and development.

Practical guides and fresh insights by email. No recycled takes.

Related dispatches

Continue reading

AI Development

Six Public Signals That an AI Model Launch Is Hours Away

Six signals that a model is shipping, five checkable and one not, from stealth tests to alias flips, with the API field for each and three launches scored.

September 21, 2026 · 9 minRead
AI Development

Grok 4.7 Costs the Same as 4.6. What Actually Changed

xAI released Grok 4.7 on September 21, 2026 at Grok 4.6's $2/$6 list price. The vendor-run benchmark table, the route price you are billed, three questions.

September 21, 2026 · 7 minRead
AI Development

Small AI Models Trained for One Job: When They Win

A 4B model trained for $1,200 cut Postgres query latency 44.7% on a standard benchmark. The trait that made it work, and a routing table for your own tasks.

September 18, 2026 · 8 minRead
AI Development

Qwen3.8-Omni-Flash: Cheaper Audio and Video AI Agents

Qwen's September 18 model takes audio and video natively with 1M context and, by Qwen's own method, cuts per-hour audio cost 98% versus its predecessor.

September 18, 2026 · 9 minRead
AI Development

Computer-Use Agents: Microsoft vs Anthropic vs Google

Microsoft GA, Anthropic public beta, and Google Gemini preview — OSWorld scores now 78% across frontier models above the ~72% human baseline. Routing guide.

May 22, 2026 · 16 minRead
AI Development

Agent Computer Use: Enterprise Automation Playbook

Enterprise playbook for deploying computer-use agents — a 40-point guardrails checklist spanning identity, audit, action boundaries, failures, and compliance.

May 22, 2026 · 17 minRead