For Ollama Qwen, run ollama run qwen3.5:9b on 8 to 16GB of memory, ollama run qwen3.8:27b on 24GB, and ollama run qwen3.6:35b-a3b on 32GB or more. All three are current Qwen releases with tools, vision, and thinking badges and a 256K context window on the Ollama library, checked on 2026-09-30 (Source: Ollama library).

Tag size is only part of the fit. The KV cache for your context window sits on top of the weights, and Ollama picks a default context from your VRAM: 4k below 24 GiB, 32k from 24 to 48 GiB, and 256k at 48 GiB or more (Source: Ollama context docs). The tables below map each Qwen family and memory tier, with the math. This guide was checked against Ollama v0.35.0, the latest stable release, published 2026-09-28 (Source: Ollama v0.35.0 release).

Key takeaways:

  • Default picks: qwen3.5:9b (6.6GB), qwen3.8:27b (18GB), and qwen3.6:35b-a3b (23GB) are the current general-purpose Qwen tags for 8, 24, and 32GB machines.
  • Context costs less on new Qwen: Qwen3.5 and later cache keys and values in only one layer out of four, so qwen3.5:9b needs 32 KiB per token against 144 KiB for qwen3:8b.
  • Thinking is on by default for qwen3.5, qwen3.6, and qwen3.8: turn it off with /set nothink in a chat or "think": false in the API.

Every Qwen model on Ollama

Ollama's library lists 18 Qwen entries; these are the ones that matter. Each tag, size, and context value comes from the model's tags page on ollama.com; the minimum Ollama version comes from each model's registry config (Source: Ollama library).

FamilyDefault tag (latest)Size on diskMax contextCapabilitiesOther local sizes
qwen3.8qwen3.8:27b18GB256Ktools, vision, thinking27B only; needs Ollama 0.32.12+
qwen3.6qwen3.6:35b-a3b23GB256Ktools, vision, thinking27b 18GB; needs 0.30.0+
qwen3.5qwen3.5:9b6.6GB256Ktools, vision, thinking0.8b 1.0GB, 2b 2.7GB, 4b 3.4GB, 27b 17GB, 35b 24GB, 122b 81GB
qwen3qwen3:8b5.2GB40Ktools, thinking0.6b to 235b; 4b, 30b, 235b default to 256K
qwen2.5qwen2.5:7b4.7GB32Ktools0.5b to 72b
qwen3-vlqwen3-vl:8b6.1GB256Ktools, vision, thinking2b to 235b
qwen2.5vlqwen2.5vl:7b6.0GB125Kvision3b, 32b, 72b
qwen3.8-flash-nextnone105GB smallest256Ktools, vision, thinking125b-a6b only

The qwen3.8, qwen3.6, and qwen3.5 downloads each include an Apache 2.0 license layer. None of the Qwen tags pages lists a cloud tag, so every row above runs on your own hardware (Source: Ollama library).

Searching for Qwen 4 on Ollama leads to qwen3.8-flash-next, which its model page calls an experimental preview of the architecture that will underpin Qwen4: 125B total parameters, 6B active per token, and a smallest tag of 105GB (Source: Ollama library).

Which Qwen to run for your memory

Pick the row that matches your GPU VRAM, or the memory left over on an Apple Silicon Mac after macOS and your apps. The last column adds the tag size to the f16 KV cache at the suggested context, using the per-token figures in the next section (Sources: Ollama library, Hugging Face Qwen configs).

MemoryRun thisSize on diskSet context toKV cache (f16)Weights plus KV
4 to 6GBqwen3.5:4b3.4GB32K1.1GBabout 4.5GB
8GBqwen3.5:9b6.6GB16K (32K is tight)0.5GB (1.1GB)about 7.1GB (7.7GB)
12 to 16GBqwen3.5:9b or qwen3.5:9b-q8_06.6GB or 11GB64K2.1GBabout 8.7GB or 12.8GB
24GBqwen3.8:27b18GB32K to 64K2.1 to 4.3GBabout 19.9 to 22.0GB
32GBqwen3.6:35b-a3b23GB64K to 128K1.3 to 2.7GBabout 24.0 to 25.3GB
48GBqwen3.8:27b18GB256K17.2GBabout 34.9GB
64GBqwen3.8:27b-q8_030GB256K17.2GBabout 47.2GB
96GB and upqwen3.5:122b or qwen3.8-flash-next:125b-a6b-nvfp481GB or 105GBas needednot calculated herecheck ollama ps

Inference: leave 1 to 2GB above the last column for Ollama's compute buffers and the display. That is why qwen3.6:35b-a3b sits in the 32GB row even though its 23GB file looks close to a 24GB card; at 24GB it runs, but expect part of it on the CPU.

The 48GB row is also the point where Ollama's own default reaches 256K, so on that hardware qwen3.8:27b gets its full window without any setting (Source: Ollama context docs). To run the same math for any other model or quant, use the LLM VRAM calculator.

Why Qwen3.5 and newer need less KV cache

Qwen3.5, Qwen3.6, and Qwen3.8 use a hybrid attention design in which three of every four layers are Gated DeltaNet linear attention and only every fourth layer is full attention. The linear layers keep a fixed-size state that does not grow with the conversation, so only the full-attention layers add to the KV cache as context grows (Source: Hugging Face Qwen configs).

The per-token cost is 2 (keys and values) x full-attention layers x KV heads x head dimension x 2 bytes at f16. For the 27B models, the config lists 64 layers with a full-attention interval of 4, so 16 layers cache, each with 4 KV heads of 256 dimensions: 2 x 16 x 4 x 256 x 2 = 65,536 bytes, or 64 KiB per token. At 32,768 tokens that is about 2.1GB (Source: Hugging Face Qwen configs).

TagsFull-attention layersKV heads x head dimKV per token (f16)At 32KAt 128K
qwen3.8:27b, qwen3.6:27b, qwen3.5:27b16 of 644 x 25664 KiB2.1GB8.6GB
qwen3.6:35b-a3b, qwen3.5:35b-a3b10 of 402 x 25620 KiB0.7GB2.7GB
qwen3.5:9b, qwen3.5:4b8 of 324 x 25632 KiB1.1GB4.3GB
qwen3:8b, qwen3-vl:8b36 of 368 x 128144 KiB4.8GB19.3GB (qwen3-vl only)
qwen2.5:7b28 of 284 x 12856 KiB1.9GBbeyond its 32K limit

The older rows come from the Qwen3-8B, Qwen3-VL-8B, and Qwen2.5-7B configs, and the 35B row from the Qwen3.6-35B-A3B config. So qwen3.5:9b holds 32K of context in 1.1GB while qwen3:8b needs 4.8GB, which makes the newer model the better 8GB pick despite its larger file. Our KV cache size guide covers the formula for other architectures.

How to pull a specific Qwen quant tag in Ollama

Every default Qwen tag is a 4-bit Q4_K_M build; the registry config for qwen3.8:27b, qwen3.6:35b-a3b, and qwen3.5:9b all report file_type Q4_K_M (Source: Ollama library). To get a different quant, name the full tag from the tags page:

ollama pull qwen3.5:9b-q8_0      # 11GB, 8-bit
ollama pull qwen3.8:27b-q8_0     # 30GB, 8-bit
ollama pull qwen3.8:27b-bf16     # 56GB, unquantized
ollama pull qwen3.8:27b-mlx      # 18GB, Apple Silicon MLX build

The tags page marks MLX builds, including the nvfp4 and mxfp8 variants, with an MLX label; those target Apple Silicon (Source: Ollama library). Tags with mtp in the name, such as qwen3.8:27b-mtp-q4_K_M, add multi-token prediction for speculative decoding, which has had its own regression reports.

The two small qwen3.5 sizes differ from the rest: qwen3.5:2b and qwen3.5:0.8b point at their q8_0 builds by default, which is why 2b is 2.7GB while 2b-q4_K_M is 1.9GB (Source: Ollama library). For whether the step from 4-bit to 5-bit or 8-bit is worth the memory, see our Q4_K_M vs Q5_K_M comparison.

qwen3.8 needs Ollama 0.32.12 or newer and qwen3.6 needs 0.30.0, so update to v0.35.0 before pulling either (Sources: Ollama library, Ollama v0.35.0 release).

How to set the context length for Qwen in Ollama

Qwen models advertise 256K, but Ollama loads them with far less unless you raise it. The documented default is 4k tokens below 24 GiB of VRAM, and Ollama recommends at least 64000 tokens for agents, web search, and coding tools (Source: Ollama context docs). The FAQ page still says the default is 4096 tokens without the VRAM tiers, so treat the context-length page as current (Source: Ollama FAQ).

You can set it in three places. For the server, which applies to every client:

OLLAMA_CONTEXT_LENGTH=32768 ollama serve

Inside an ollama run qwen3.5:9b session, type /set parameter num_ctx 32768. Through the API, pass "options": {"num_ctx": 32768} in the request body (Source: Ollama FAQ).

Then confirm it took. ollama ps shows a CONTEXT column and a PROCESSOR column; you want your number in the first and 100% GPU in the second (Source: Ollama context docs). If the split shows CPU, lower the context or quantize the cache: OLLAMA_KV_CACHE_TYPE=q8_0 uses about half the memory of the default f16 and only works when Flash Attention is active. Parallel requests multiply the cache too, because RAM scales with OLLAMA_NUM_PARALLEL times the context length (Source: Ollama FAQ). The full list of environment variables is in our Ollama commands cheat sheet.

How to turn Qwen thinking on and off in Ollama

Thinking-capable Qwen models return their reasoning in a separate thinking field, and the think request field controls it: true asks for thinking, false asks for none, null uses the model default, and a string picks a named level (Source: Ollama thinking docs). Ollama's v0.35.0 renderer setup starts qwen3.5 (which qwen3.6 also uses) and qwen3.8 with thinking on.

Pick the method that matches how you run the model:

  • Interactive chat: type /set nothink to turn thinking off and /set think to turn it back on, both listed in the session's help text in cmd/interactive.go.
  • One-shot CLI: ollama run qwen3.5:9b --think=false "summarize this" skips it, and --hidethinking keeps the reasoning but prints only the answer, per Ollama's thinking announcement.
  • API: send "think": false in the /api/chat or /api/generate body.

qwen3.8 is different from the older Qwen models, which only take true or false. In the v0.35.0 source, qwen3.8 accepts false, "low", "medium", and "xhigh", and uses "xhigh" when think is not set (qwen35.go at v0.35.0). Run curl http://localhost:11434/api/show -d '{"model":"qwen3.8"}' to see the values your install supports (Source: Ollama thinking docs).

Stick to those values: an open issue reports that "high" and "max" were silently accepted on 0.34.x and ran the default, then medium (ollama#18632). Inference: on a small card, try "low" before switching thinking off.

Qwen coder tags on Ollama

Qwen also ships coding-tuned tags, and our best Ollama model for coding guide ranks them against Gemma, gpt-oss, and Devstral by VRAM and agent tool. The tags themselves, from the library on 2026-09-30, are these (Source: Ollama library):

  • Current Qwen coding tags: qwen3.6:27b-coding (18GB) and qwen3.6:35b-a3b-coding (23GB), both 256K with tools, vision, and thinking.
  • Dedicated coder models: qwen3-coder:30b (19GB, 256K, tools) and qwen3-coder-next (52GB, 256K, tools).
  • Autocomplete-sized: qwen2.5-coder:1.5b (986MB) and qwen2.5-coder:7b (4.7GB), both limited to 32K.

Known issues when running Qwen on Ollama

Most Qwen problems on Ollama trace to three things: tool-call parsing, long agent loops, and memory on Apple Silicon. These were open on GitHub on 2026-09-30:

  • Long tool loops on qwen3.8: ollama#17778 reports a 500 error, "no user query found in messages", after many tool calls at around 205K context. It has 47 comments.
  • Tool calls on Qwen3 via /api/chat: ollama#14601 reports malformed tool definitions when tools are passed through the tools parameter.
  • Long file writes from coder models: ollama#18563 reports the qwen3coder parser rejecting long file-write tool calls.
  • Memory growth on MLX: ollama#18620 reports qwen3.6:27b-mlx holding about 0.43 GiB more after each request that ends in a tool call, on a 32GB Mac.

Inference: when an agent fails on Qwen, check ollama ps for a context above the 4k default first, then retry the GGUF tag instead of the -mlx one (Source: Ollama context docs).

FAQ

Is Qwen available in Ollama?

Yes. The Ollama library hosts every major Qwen generation, from qwen (Qwen 1.5) through qwen2.5, qwen3, qwen3.5, qwen3.6, and qwen3.8, plus coder and vision variants. Run ollama run qwen3.5 to start the default 9B model, or add a size tag such as qwen3.8:27b for a larger one.

How to install Qwen on Ollama?

Install Ollama, then run ollama pull qwen3.5:9b to download the model and ollama run qwen3.5:9b to chat with it. Choose the tag for your memory from the table above. On 24GB or more, ollama run qwen3.8:27b gets the newest release. Update Ollama first, because qwen3.8 requires 0.32.12 or later.

Is Ollama Qwen 3.5 free?

Yes. Ollama is free to run locally, and the qwen3.5 weights ship with an Apache 2.0 license, which allows commercial use. Running it costs only your own hardware and power. The Qwen tags pages list no cloud tags, so every qwen3.5 tag downloads and runs on your machine with no account or usage fee.

What are the limitations of Qwen?

On Ollama, the main limits are memory and defaults rather than the model. Below 24 GiB of VRAM Ollama starts Qwen at a 4k context, far below its 256K maximum. Thinking is on by default and slows short answers, and several tool-calling bugs remain open for Qwen models in long agent sessions.

Is Ollama like chatgpt?

Partly. Ollama gives you a chat window and an API like ChatGPT, but the model runs on your own computer instead of a cloud service. Answers stay local and work offline once the model is downloaded. The trade-off is that speed and model size depend entirely on your GPU or unified memory.

References