AI DevelopmentMethodology7 min readPublished September 17, 2026

60 models · rolling 30-day downloads · collected September 18, 2026

Which Open AI Models People Download: September 2026 Data

The 60 most-downloaded open text and vision-language models on Hugging Face, by 30-day downloads collected September 18, 2026. Qwen holds 27 rows; the caveats.

DA
Digital Applied Team
Research and practical guidance
Editorial dateSeptember 17, 2026
SourceHugging Face Hub API, our own query

Ask which open AI models people actually use and most answers are opinions. Download counts are not a perfect answer, but they are a measured one. On September 18, 2026 we queried the Hugging Face Hub, the main public host for open model weights, for the most-downloaded models in its three language and vision-language categories and kept the top 60 by downloads over the previous 30 days. Qwen repositories hold 27 of the 60 rows. Qwen3-0.6B, a model small enough to run on a phone, leads at 22.5 million.

This page is for anyone choosing an open model who wants to know what the rest of the ecosystem is pulling, and for writers who need a dated, reproducible adoption table rather than a vibe. It explains what a Hugging Face download does and does not count, prints all 60 rows, sums them by model family, and says plainly where downloads stop being evidence. The query is stated in the methodology so anyone can repeat it.

Key takeaways
  1. 01
    Qwen dominates the top 60.27 rows are Qwen's own repositories and 36 are Qwen-based once community repacks are counted. Those 36 rows total 240 million downloads, about seven times the next family, Gemma.
  2. 02
    Small models win on raw downloads.Nine of the 17 non-fixture models in the top 20 are 9B parameters or fewer, and four more are mixture-of-experts models with 3 to 4B active. Downloads scale with how many machines a model fits on.
  3. 03
    Test fixtures are in the top 15, and they are flagged.GPT-2, a tiny Qwen2 test model and OPT-125m are downloaded by CI pipelines, not users. We show them because the raw ranking includes them, and mark them so nobody quotes them as adoption.
  4. 04
    Seventeen rows are repacks; 44 are Apache 2.0.GGUF, FP8, NVFP4, AWQ and MLX conversions of other models fill 17 of 60 rows. The licence mix is 44 Apache 2.0, 6 MIT, 4 custom, 3 Meta or Google model licences, 1 OpenRAIL and 2 unstated.

01DefinitionsWhat a download counts

The Hugging Face Hub reports a "downloads" number per model repository. In the Hub API this figure is a rolling count over the last 30 days, which is why it moves every day and why the date of collection matters. It counts file fetches from the repository, which means a person downloading a model once, a server pulling weights on every fresh deployment, and a continuous-integration job fetching a tiny model to run a test all add to the same number. It does not count a model that was downloaded once and then copied around a company, and it does not count use through an API. The Hub's API documentation describes the endpoint and its sort options.

Three consequences follow, and they shape how to read the table. Small models rank high because they fit in more places and are cheap to re-download. Test fixtures rank high because CI runs many times a day. And a quantised repack of a popular model counts separately from the original, so a model family's real reach is the sum of its rows, not its best one. We flag the fixtures with a dagger, keep every row otherwise, and sum by family in section 03.

02DatasetThe top 60, September 2026

Rank is by rolling 30-day downloads at collection time. Model is the repository name, organisation included, exactly as the Hub returns it. Licence is the tag the repository declares; "Other (custom)" is the Hub's own category for a licence file that is not a standard one. A dagger marks a repository we consider a test fixture: a model whose downloads come mainly from automated test suites.

Source: Hugging Face Hub API, /api/models sorted by downloads, pipeline tags text-generation, image-text-to-text and any-to-any, queried September 18, 2026 09:28 UTC. † = test fixture, shown but not adoption.
#Model repository30-day downloadsLicence
1Qwen/Qwen3-0.6B22,498,727Apache 2.0
2Qwen/Qwen3-VL-8B-Instruct19,098,599Apache 2.0
3openai-community/gpt2 †15,439,333MIT
4trl-internal-testing/tiny-Qwen2ForCausalLM-2.5 †14,140,672Not stated
5Qwen/Qwen3-8B12,988,756Apache 2.0
6unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF12,752,716Apache 2.0
7Qwen/Qwen3.6-35B-A3B-FP810,210,602Apache 2.0
8google/gemma-4-26B-A4B-it9,794,984Apache 2.0
9Qwen/Qwen2.5-7B-Instruct9,714,062Apache 2.0
10Qwen/Qwen3.5-9B9,285,215Apache 2.0
11google/gemma-4-31B-it9,045,393Apache 2.0
12nvidia/Qwen3.6-35B-A3B-NVFP48,594,317Apache 2.0
13Qwen/Qwen2.5-0.5B-Instruct8,490,193Apache 2.0
14Qwen/Qwen3.8-27B-FP87,603,458Apache 2.0
15facebook/opt-125m †7,486,472Other (custom)
16Qwen/Qwen3.8-27B7,456,257Apache 2.0
17farbodtavakkoli/OTel-2.0-LLM-31B-IT7,407,454Apache 2.0
18Qwen/Qwen3-4B7,242,404Apache 2.0
19Qwen/Qwen2.5-1.5B-Instruct7,184,297Apache 2.0
20Qwen/Qwen3.5-4B6,958,404Apache 2.0
21meta-llama/Llama-3.2-1B-Instruct6,921,311Llama 3.2
22Qwen/Qwen2.5-VL-7B-Instruct6,851,431Apache 2.0
23openai/gpt-oss-20b6,675,065Apache 2.0
24Qwen/Qwen3.6-27B-FP86,150,001Apache 2.0
25meta-llama/Llama-3.1-8B-Instruct5,934,139Llama 3.1
26ornith-ai/Ornith-1.5-9B-GGUF5,532,094MIT
27openai/gpt-oss-120b5,185,534Apache 2.0
28Qwen/Qwen2.5-3B-Instruct5,114,042Other (custom)
29Qwen/Qwen3-32B4,958,092Apache 2.0
30lmstudio-community/Qwen3.8-27B-MLX-4bit4,873,208Apache 2.0
31dphn/dolphin-2.9.1-yi-1.5-34b4,820,618Apache 2.0
32Qwen/Qwen3.5-2B4,811,106Apache 2.0
33lmstudio-community/Qwen3.8-27B-MLX-8bit4,665,399Apache 2.0
34lmstudio-community/Qwen3.8-27B-MLX-6bit4,620,470Apache 2.0
35lmstudio-community/Qwen3.8-27B-MLX-5bit4,591,269Apache 2.0
36ornith-ai/Ornith-1.5-35B-A3B-GGUF4,492,820MIT
37google/gemma-4-E4B-it4,473,231Apache 2.0
38deepseek-ai/DeepSeek-V4-Flash-07314,394,898MIT
39Qwen/Qwen-72B4,014,339Other (custom)
40Qwen/Qwen3-4B-Instruct-25073,992,032Apache 2.0
41Qwen/Qwen3-VL-4B-Instruct3,798,548Apache 2.0
42Qwen/Qwen3.6-27B3,777,499Apache 2.0
43Qwen/Qwen3-1.7B3,756,550Apache 2.0
44Qwen/Qwen2.5-7B-Instruct-AWQ3,507,773Apache 2.0
45google/gemma-4-E2B-it3,500,771Apache 2.0
46nvidia/NVIDIA-Nemotron-3-Nano-4B-BF163,484,079Other (custom)
47EleutherAI/pythia-160m3,469,135Apache 2.0
48Qwen/Qwen3.6-35B-A3B3,452,957Apache 2.0
49ornith-ai/Ornith-1.0-9B-GGUF3,443,588MIT
50cdiamond/Qwen3.8-27B-iMatrix-NVFP4-MTP-GGUF3,399,484Apache 2.0
51RadixArk/Kimi-K3-DSpark3,338,184Not stated
52Qwen/Qwen3-VL-2B-Instruct3,078,197Apache 2.0
53google/gemma-3-1b-it3,018,334Gemma
54microsoft/Florence-2-base3,010,170MIT
55huihui-ai/Huihui-Qwen3.8-27B-abliterated-GGUF2,859,860Apache 2.0
56datalab-to/chandra-ocr-22,836,048OpenRAIL
57google/gemma-4-12B-it2,802,172Apache 2.0
58Qwen/Qwen2.5-Coder-7B-Instruct2,672,084Apache 2.0
59Qwen/Qwen3-14B-AWQ2,555,771Apache 2.0
60JonathanColetti/Qwen3.8-27B-Uncensored-GGUF2,545,055Apache 2.0

03FamiliesDownloads by model family

Summing rows by the model they derive from, and leaving out the three fixtures, gives the chart below. We assigned a family by the repository name and its declared base model, so a community GGUF of Qwen3.8-27B counts as Qwen and Nvidia's NVFP4 build of Qwen3.6-35B-A3B counts as Qwen. The 10 largest families are shown; the remaining three rows, Kimi K3, Florence-2 and Chandra OCR, are each a single repository under 3.4 million.

30-day downloads by model family, top 60 excluding fixtures

Our sum of the table above, Hugging Face Hub API, September 18, 2026. Family assigned by repository name and declared base model.
Qwen36 rows
240.1M
Gemma6 rows
32.6M
Ornith3 rows
13.5M
Llama2 rows
12.9M
gpt-oss2 rows
11.9M
OTel1 row
7.4M
Yi1 row
4.8M
DeepSeek1 row
4.4M
Nemotron1 row
3.5M
Pythia1 row
3.5M

Two families deserve a note. Ornith's three GGUF repositories sum to 13.5 million, which puts a name most readers will not know above Llama and gpt-oss; we have not investigated the source of those downloads and record the number as the Hub reports it. And Llama has two rows in the top 60, both from the 3.x generation released in 2024, which fits the shift toward Qwen and DeepSeek described in our first-half 2026 open-weight retrospective.

04FindingsFour things the table shows

Share
Rows that are Qwen's own repositories
27of 60

Google has 6, lmstudio-community 4, Ornith 3, and Nvidia, Meta and OpenAI 2 each. No other organisation has more than one row.

By organisation
Size
Non-fixture top-20 models at 9B parameters or fewer
9of 17

Qwen3-0.6B, Qwen3-VL-8B, Qwen3-8B, Qwen2.5-7B, Qwen3.5-9B, Qwen2.5-0.5B, Qwen3-4B, Qwen2.5-1.5B and Qwen3.5-4B. Four more are MoE models with 3 to 4B active.

By parameter count
Repacks
Quantised or converted copies of another model
17of 60

Four are lmstudio-community MLX builds of Qwen3.8-27B alone, at 4, 5, 6 and 8 bits, totalling 18.8 million. The FP8 build of Qwen3.8-27B out-downloads the original, 7.6 to 7.5 million.

By repository type
Recency
Repositories created in 2026
30of 60

Half the table is under nine months old. The oldest non-fixture rows are Pythia-160m and Qwen-72B from 2023; Qwen2.5 models from late 2024 and January 2025 still hold seven rows.

By creation date

The repack finding is the practical one. If you want to know which size and format of a model people actually run, the quantised rows answer it more directly than the originals: an FP8 or 4-bit file is what gets loaded onto a real machine. Our self-hosting guide for open coding models pairs those formats with hardware, and this week's comparison of SSD streaming and ternary weights covers two newer ways to make a large model fit; the Edge0 model from that post was the Hub's top trending repository on the day we collected, with 37,131 downloads and 3,349 likes.

05LimitsWhat it cannot tell you

Downloads are not users, and users are not production. A download count says how many times files were fetched from one host in 30 days. It cannot distinguish a person from a script, a trial from a deployment, or ten thousand developers from one company's build farm. It excludes every model served through an API, every mirror such as ModelScope, and every weight file that was fetched once and then cached inside an organisation. A model that is downloaded by a few hundred companies and run at scale can rank below a model that is fetched by a hundred thousand hobbyists once each.

It also cannot tell you where use is. We have not computed a country share and do not print one: a repository's organisation says where a model was made, not where it is run. Treat the table as a measure of distribution, read it alongside benchmark and price data, and check the licence column before assuming any row is usable in a product. If you want help choosing an open model for a specific workload, our AI transformation service includes that evaluation.

The fixtures, named

openai-community/gpt2 (15.4 million), trl-internal-testing/ tiny-Qwen2ForCausalLM-2.5 (14.1 million) and facebook/opt-125m (7.5 million) are downloaded by test suites for training and inference libraries. Together they are 37 million of the table's 381 million downloads. EleutherAI/pythia-160m at row 47 is a research model that is also common in tests; we left it unmarked because we cannot separate its uses.

06How we collected itMethodology

Methodology

This is our own measurement from a public API. Nothing here comes from a vendor's report or a third-party ranking.

Query
Hugging Face Hub API, GET /api/models, with pipeline_tag set to each of text-generation, image-text-to-text and any-to-any in turn, sort set to downloads, descending, top 100 per tag. The three lists were merged and the top 60 by downloads kept. The trending list was collected separately with the same endpoint's trending sort.
What the count is
The Hub's downloads field, a rolling count of the previous 30 days at query time. Likes, creation date, licence tag and declared base model were recorded from the same response.
Scope
Language and vision-language models only. Image, audio, video and embedding models are excluded by the tag filter. Models without one of the three tags, however popular, do not appear.
Fixture rule
A row is marked as a fixture when the repository is a known test model for a library, its organisation is an internal-testing account, or it is a 2022-era baseline whose current downloads are implausible as human use. Three rows meet the rule. Fixtures are shown in the table and excluded from the family chart.
As-of date
Collected September 18, 2026 at 09:28 UTC. This page is dated to the editorial day before; the collection date is stated here and in the dataset card. The next pull will be a new dated snapshot, not a revision of this one.
Known limitations
One host, one 30-day window, no de-duplication of repeated fetches, no API usage, no mirrors. Family assignment by name and base-model tag can misclassify a fine-tune that does not declare its base. Parameter counts in the findings are read from repository names.

07Next stepOne family, many small files

Put it into practice

Check the repack rows before you pick a size and format

If you are choosing an open model this month, the table's most useful signal is not the winner but the shape: small models and quantised copies of mid-sized ones are what the ecosystem loads. Find the family you are considering, look at which of its formats people fetch, and test that file on your hardware. Come back next month; the next pull will be a new dated row in the same series.

Digital Applied

Choose an open model on evidence, not on the leaderboard of the week.

We evaluate open models against your own prompts, hardware and licence needs, and keep the adoption data behind pages like this one so the shortlist reflects what teams actually run.

Open-model shortlistLicence reviewLocal deployment
Your next project

Start with the format people load

  • Find your family in the table
  • Pick the repack size that fits your machine
  • Test on twenty of your own prompts
Questions and answers

Applying this post

Because software test suites download it. GPT-2 is small, permissively licensed and supported everywhere, so libraries use it as a fixture in automated tests that run many times a day. The same is true of the tiny Qwen2 test model and OPT-125m. We mark all three with a dagger and leave them out of the family totals.
Digital Applied newsletter

Deep dives on AI, marketing and development.

Practical guides and fresh insights by email. No recycled takes.

Related dispatches

Continue reading

AI Development

Run a 35B AI Model on a 24GB Mac: Two New Ways Compared

Edge0 streams a 35B mixture-of-experts model from SSD in under 3 GiB of RAM. Bonsai 2 shrinks a 27B model to 5.9 GB. Which to pick, and what to test first.

September 17, 2026 · 7 minRead
AI Development

Best Open-Weight Coding Models to Self-Host in 2026

Match 2026's best open-weight coding models to your hardware: Qwen3-Coder-Next, Devstral 2, GLM-5.2 and DeepSeek V4 by VRAM, SWE-bench score and real speed.

June 29, 2026 · 14 minRead
AI Development

Is Your Local AI Really Offline? A Practical Test Guide

Test whether a local AI workflow depends on cloud models, search, embeddings or downloads, and record what works offline without overstating privacy guarantees.

September 13, 2026 · 5 minRead
AI Development

Kimi K3 Open Weights: A July 27 Readiness Checklist

Kimi K3's open weights are promised by July 27, not shipped — use the 10-day window to prep hosting for a 2.8T model and check the license before you commit.

July 17, 2026 · 12 minRead
AI Development

AI PCs and NPUs in 2026: Can They Really Run Local AI?

AI PCs clear the 40+ TOPS Copilot+ bar, but can NPUs actually run local LLMs? A 2026 buyer's guide to NPU vs GPU, Phi Silica and what really matters.

June 29, 2026 · 12 minRead
AI Development

Small Language Models for On-Device Agents in 2026

A 3-9B small language model on your laptop can handle most agentic loop steps faster and cheaper than a frontier cloud model. Here's how to build SLM-first.

June 29, 2026 · 12 minRead