For Ollama Qwen, run ollama run qwen3.5:9b on 8 to 16GB of memory, ollama run qwen3.8:27b on 24GB, and ollama run qwen3.6:35b-a3b on 32GB or more. All three are current Qwen releases with tools, vision, and thinking badges and a 256K context window on the Ollama library, checked on 2026-09-30 (Source: Ollama library).
Tag size is only part of the fit. The KV cache for your context window sits on top of the weights, and Ollama picks a default context from your VRAM: 4k below 24 GiB, 32k from 24 to 48 GiB, and 256k at 48 GiB or more (Source: Ollama context docs). The tables below map each Qwen family and memory tier, with the math. This guide was checked against Ollama v0.35.0, the latest stable release, published 2026-09-28 (Source: Ollama v0.35.0 release).
Key takeaways:
- Default picks:
qwen3.5:9b(6.6GB),qwen3.8:27b(18GB), andqwen3.6:35b-a3b(23GB) are the current general-purpose Qwen tags for 8, 24, and 32GB machines. - Context costs less on new Qwen: Qwen3.5 and later cache keys and values in only one layer out of four, so
qwen3.5:9bneeds 32 KiB per token against 144 KiB forqwen3:8b. - Thinking is on by default for
qwen3.5,qwen3.6, andqwen3.8: turn it off with/set nothinkin a chat or"think": falsein the API.
Every Qwen model on Ollama
Ollama's library lists 18 Qwen entries; these are the ones that matter. Each tag, size, and context value comes from the model's tags page on ollama.com; the minimum Ollama version comes from each model's registry config (Source: Ollama library).
| Family | Default tag (latest) | Size on disk | Max context | Capabilities | Other local sizes |
|---|---|---|---|---|---|
qwen3.8 | qwen3.8:27b | 18GB | 256K | tools, vision, thinking | 27B only; needs Ollama 0.32.12+ |
qwen3.6 | qwen3.6:35b-a3b | 23GB | 256K | tools, vision, thinking | 27b 18GB; needs 0.30.0+ |
qwen3.5 | qwen3.5:9b | 6.6GB | 256K | tools, vision, thinking | 0.8b 1.0GB, 2b 2.7GB, 4b 3.4GB, 27b 17GB, 35b 24GB, 122b 81GB |
qwen3 | qwen3:8b | 5.2GB | 40K | tools, thinking | 0.6b to 235b; 4b, 30b, 235b default to 256K |
qwen2.5 | qwen2.5:7b | 4.7GB | 32K | tools | 0.5b to 72b |
qwen3-vl | qwen3-vl:8b | 6.1GB | 256K | tools, vision, thinking | 2b to 235b |
qwen2.5vl | qwen2.5vl:7b | 6.0GB | 125K | vision | 3b, 32b, 72b |
qwen3.8-flash-next | none | 105GB smallest | 256K | tools, vision, thinking | 125b-a6b only |
The qwen3.8, qwen3.6, and qwen3.5 downloads each include an Apache 2.0 license layer. None of the Qwen tags pages lists a cloud tag, so every row above runs on your own hardware (Source: Ollama library).
Searching for Qwen 4 on Ollama leads to qwen3.8-flash-next, which its model page calls an experimental preview of the architecture that will underpin Qwen4: 125B total parameters, 6B active per token, and a smallest tag of 105GB (Source: Ollama library).
Which Qwen to run for your memory
Pick the row that matches your GPU VRAM, or the memory left over on an Apple Silicon Mac after macOS and your apps. The last column adds the tag size to the f16 KV cache at the suggested context, using the per-token figures in the next section (Sources: Ollama library, Hugging Face Qwen configs).
| Memory | Run this | Size on disk | Set context to | KV cache (f16) | Weights plus KV |
|---|---|---|---|---|---|
| 4 to 6GB | qwen3.5:4b | 3.4GB | 32K | 1.1GB | about 4.5GB |
| 8GB | qwen3.5:9b | 6.6GB | 16K (32K is tight) | 0.5GB (1.1GB) | about 7.1GB (7.7GB) |
| 12 to 16GB | qwen3.5:9b or qwen3.5:9b-q8_0 | 6.6GB or 11GB | 64K | 2.1GB | about 8.7GB or 12.8GB |
| 24GB | qwen3.8:27b | 18GB | 32K to 64K | 2.1 to 4.3GB | about 19.9 to 22.0GB |
| 32GB | qwen3.6:35b-a3b | 23GB | 64K to 128K | 1.3 to 2.7GB | about 24.0 to 25.3GB |
| 48GB | qwen3.8:27b | 18GB | 256K | 17.2GB | about 34.9GB |
| 64GB | qwen3.8:27b-q8_0 | 30GB | 256K | 17.2GB | about 47.2GB |
| 96GB and up | qwen3.5:122b or qwen3.8-flash-next:125b-a6b-nvfp4 | 81GB or 105GB | as needed | not calculated here | check ollama ps |
Inference: leave 1 to 2GB above the last column for Ollama's compute buffers and the display. That is why qwen3.6:35b-a3b sits in the 32GB row even though its 23GB file looks close to a 24GB card; at 24GB it runs, but expect part of it on the CPU.
The 48GB row is also the point where Ollama's own default reaches 256K, so on that hardware qwen3.8:27b gets its full window without any setting (Source: Ollama context docs). To run the same math for any other model or quant, use the LLM VRAM calculator.
Why Qwen3.5 and newer need less KV cache
Qwen3.5, Qwen3.6, and Qwen3.8 use a hybrid attention design in which three of every four layers are Gated DeltaNet linear attention and only every fourth layer is full attention. The linear layers keep a fixed-size state that does not grow with the conversation, so only the full-attention layers add to the KV cache as context grows (Source: Hugging Face Qwen configs).
The per-token cost is 2 (keys and values) x full-attention layers x KV heads x head dimension x 2 bytes at f16. For the 27B models, the config lists 64 layers with a full-attention interval of 4, so 16 layers cache, each with 4 KV heads of 256 dimensions: 2 x 16 x 4 x 256 x 2 = 65,536 bytes, or 64 KiB per token. At 32,768 tokens that is about 2.1GB (Source: Hugging Face Qwen configs).
| Tags | Full-attention layers | KV heads x head dim | KV per token (f16) | At 32K | At 128K |
|---|---|---|---|---|---|
qwen3.8:27b, qwen3.6:27b, qwen3.5:27b | 16 of 64 | 4 x 256 | 64 KiB | 2.1GB | 8.6GB |
qwen3.6:35b-a3b, qwen3.5:35b-a3b | 10 of 40 | 2 x 256 | 20 KiB | 0.7GB | 2.7GB |
qwen3.5:9b, qwen3.5:4b | 8 of 32 | 4 x 256 | 32 KiB | 1.1GB | 4.3GB |
qwen3:8b, qwen3-vl:8b | 36 of 36 | 8 x 128 | 144 KiB | 4.8GB | 19.3GB (qwen3-vl only) |
qwen2.5:7b | 28 of 28 | 4 x 128 | 56 KiB | 1.9GB | beyond its 32K limit |
The older rows come from the Qwen3-8B, Qwen3-VL-8B, and Qwen2.5-7B configs, and the 35B row from the Qwen3.6-35B-A3B config. So qwen3.5:9b holds 32K of context in 1.1GB while qwen3:8b needs 4.8GB, which makes the newer model the better 8GB pick despite its larger file. Our KV cache size guide covers the formula for other architectures.
How to pull a specific Qwen quant tag in Ollama
Every default Qwen tag is a 4-bit Q4_K_M build; the registry config for qwen3.8:27b, qwen3.6:35b-a3b, and qwen3.5:9b all report file_type Q4_K_M (Source: Ollama library). To get a different quant, name the full tag from the tags page:
ollama pull qwen3.5:9b-q8_0 # 11GB, 8-bit
ollama pull qwen3.8:27b-q8_0 # 30GB, 8-bit
ollama pull qwen3.8:27b-bf16 # 56GB, unquantized
ollama pull qwen3.8:27b-mlx # 18GB, Apple Silicon MLX buildThe tags page marks MLX builds, including the nvfp4 and mxfp8 variants, with an MLX label; those target Apple Silicon (Source: Ollama library). Tags with mtp in the name, such as qwen3.8:27b-mtp-q4_K_M, add multi-token prediction for speculative decoding, which has had its own regression reports.
The two small qwen3.5 sizes differ from the rest: qwen3.5:2b and qwen3.5:0.8b point at their q8_0 builds by default, which is why 2b is 2.7GB while 2b-q4_K_M is 1.9GB (Source: Ollama library). For whether the step from 4-bit to 5-bit or 8-bit is worth the memory, see our Q4_K_M vs Q5_K_M comparison.
qwen3.8 needs Ollama 0.32.12 or newer and qwen3.6 needs 0.30.0, so update to v0.35.0 before pulling either (Sources: Ollama library, Ollama v0.35.0 release).
How to set the context length for Qwen in Ollama
Qwen models advertise 256K, but Ollama loads them with far less unless you raise it. The documented default is 4k tokens below 24 GiB of VRAM, and Ollama recommends at least 64000 tokens for agents, web search, and coding tools (Source: Ollama context docs). The FAQ page still says the default is 4096 tokens without the VRAM tiers, so treat the context-length page as current (Source: Ollama FAQ).
You can set it in three places. For the server, which applies to every client:
OLLAMA_CONTEXT_LENGTH=32768 ollama serve
Inside an ollama run qwen3.5:9b session, type /set parameter num_ctx 32768. Through the API, pass "options": {"num_ctx": 32768} in the request body (Source: Ollama FAQ).
Then confirm it took. ollama ps shows a CONTEXT column and a PROCESSOR column; you want your number in the first and 100% GPU in the second (Source: Ollama context docs). If the split shows CPU, lower the context or quantize the cache: OLLAMA_KV_CACHE_TYPE=q8_0 uses about half the memory of the default f16 and only works when Flash Attention is active. Parallel requests multiply the cache too, because RAM scales with OLLAMA_NUM_PARALLEL times the context length (Source: Ollama FAQ). The full list of environment variables is in our Ollama commands cheat sheet.
How to turn Qwen thinking on and off in Ollama
Thinking-capable Qwen models return their reasoning in a separate thinking field, and the think request field controls it: true asks for thinking, false asks for none, null uses the model default, and a string picks a named level (Source: Ollama thinking docs). Ollama's v0.35.0 renderer setup starts qwen3.5 (which qwen3.6 also uses) and qwen3.8 with thinking on.
Pick the method that matches how you run the model:
- Interactive chat: type
/set nothinkto turn thinking off and/set thinkto turn it back on, both listed in the session's help text incmd/interactive.go. - One-shot CLI:
ollama run qwen3.5:9b --think=false "summarize this"skips it, and--hidethinkingkeeps the reasoning but prints only the answer, per Ollama's thinking announcement. - API: send
"think": falsein the/api/chator/api/generatebody.
qwen3.8 is different from the older Qwen models, which only take true or false. In the v0.35.0 source, qwen3.8 accepts false, "low", "medium", and "xhigh", and uses "xhigh" when think is not set (qwen35.go at v0.35.0). Run curl http://localhost:11434/api/show -d '{"model":"qwen3.8"}' to see the values your install supports (Source: Ollama thinking docs).
Stick to those values: an open issue reports that "high" and "max" were silently accepted on 0.34.x and ran the default, then medium (ollama#18632). Inference: on a small card, try "low" before switching thinking off.
Qwen coder tags on Ollama
Qwen also ships coding-tuned tags, and our best Ollama model for coding guide ranks them against Gemma, gpt-oss, and Devstral by VRAM and agent tool. The tags themselves, from the library on 2026-09-30, are these (Source: Ollama library):
- Current Qwen coding tags:
qwen3.6:27b-coding(18GB) andqwen3.6:35b-a3b-coding(23GB), both 256K with tools, vision, and thinking. - Dedicated coder models:
qwen3-coder:30b(19GB, 256K, tools) andqwen3-coder-next(52GB, 256K, tools). - Autocomplete-sized:
qwen2.5-coder:1.5b(986MB) andqwen2.5-coder:7b(4.7GB), both limited to 32K.
Known issues when running Qwen on Ollama
Most Qwen problems on Ollama trace to three things: tool-call parsing, long agent loops, and memory on Apple Silicon. These were open on GitHub on 2026-09-30:
- Long tool loops on
qwen3.8: ollama#17778 reports a 500 error, "no user query found in messages", after many tool calls at around 205K context. It has 47 comments. - Tool calls on Qwen3 via
/api/chat: ollama#14601 reports malformed tool definitions when tools are passed through thetoolsparameter. - Long file writes from coder models: ollama#18563 reports the
qwen3coderparser rejecting long file-write tool calls. - Memory growth on MLX: ollama#18620 reports
qwen3.6:27b-mlxholding about 0.43 GiB more after each request that ends in a tool call, on a 32GB Mac.
Inference: when an agent fails on Qwen, check ollama ps for a context above the 4k default first, then retry the GGUF tag instead of the -mlx one (Source: Ollama context docs).
FAQ
Is Qwen available in Ollama?
Yes. The Ollama library hosts every major Qwen generation, from qwen (Qwen 1.5) through qwen2.5, qwen3, qwen3.5, qwen3.6, and qwen3.8, plus coder and vision variants. Run ollama run qwen3.5 to start the default 9B model, or add a size tag such as qwen3.8:27b for a larger one.
How to install Qwen on Ollama?
Install Ollama, then run ollama pull qwen3.5:9b to download the model and ollama run qwen3.5:9b to chat with it. Choose the tag for your memory from the table above. On 24GB or more, ollama run qwen3.8:27b gets the newest release. Update Ollama first, because qwen3.8 requires 0.32.12 or later.
Is Ollama Qwen 3.5 free?
Yes. Ollama is free to run locally, and the qwen3.5 weights ship with an Apache 2.0 license, which allows commercial use. Running it costs only your own hardware and power. The Qwen tags pages list no cloud tags, so every qwen3.5 tag downloads and runs on your machine with no account or usage fee.
What are the limitations of Qwen?
On Ollama, the main limits are memory and defaults rather than the model. Below 24 GiB of VRAM Ollama starts Qwen at a 4k context, far below its 256K maximum. Thinking is on by default and slows short answers, and several tool-calling bugs remain open for Qwen models in long agent sessions.
Is Ollama like chatgpt?
Partly. Ollama gives you a chat window and an API like ChatGPT, but the model runs on your own computer instead of a cloud service. Answers stay local and work offline once the model is downloaded. The trade-off is that speed and model size depend entirely on your GPU or unified memory.
Related coverage
- Best Ollama Model for Coding: Picks by VRAM and Agent Tool
- KV Cache Size: How Much VRAM It Uses and How to Cut It
- LLM VRAM calculator
- Ollama commands cheat sheet
References
- Hugging Face Qwen configs - https://huggingface.co/Qwen/Qwen3.8-27B/blob/main/config.json
- Ollama context docs - https://docs.ollama.com/context-length
- Ollama FAQ - https://docs.ollama.com/faq
- Ollama library - https://ollama.com/search?q=qwen
- Ollama thinking docs - https://docs.ollama.com/capabilities/thinking
- Ollama v0.35.0 release - https://github.com/ollama/ollama/releases/tag/v0.35.0

