To run llama.cpp on Windows, download llama-<build>-bin-win-cuda-13.4-x64.zip plus cudart-llama-bin-win-cuda-13.4-x64.zip for an NVIDIA RTX 20-series or newer card, win-vulkan-x64.zip for AMD or Intel graphics, or win-cpu-x64.zip with no GPU. Unzip it into one folder and run .\llama-cli.exe -hf ggml-org/Qwen3.5-0.8B-GGUF in PowerShell. That command downloads a small model from Hugging Face and opens a chat in the terminal.
The asset names below come from the GitHub releases API on September 30, 2026, where the newest Windows build was b11292. Build numbers change several times a day, but the naming pattern has been stable, so match the suffix rather than the number (Source: llama.cpp releases).
Key takeaways
- NVIDIA: the CUDA zip does not include the CUDA runtime DLLs. Download the matching
cudart-llama-bin-win-cuda-*zip and extract both into the same folder. - AMD and Intel: start with the Vulkan zip. It is also what
winget install llama.cppinstalls. - No GPU: one
win-cpu-x64zip now covers old and new CPUs; there is no separate AVX2 or AVX-512 download. - Verify offload: look for
offloaded N/N layers to GPUin the startup log. If N is 0, you are running on the CPU.
This page is research synthesis from the release assets, install scripts, and issue tracker. It was not tested on a Windows machine.
Which llama.cpp Windows download matches your GPU?
llama.cpp is the C/C++ inference engine behind a lot of local LLM tooling. It runs GGUF model files on your CPU or GPU. The project publishes ten Windows zips per build. Each one bundles a different GPU backend: the code path that talks to your graphics hardware (Source: llama.cpp releases).
| Your hardware | Download this zip | Also needed |
|---|---|---|
| NVIDIA RTX 20-series or newer, driver 580+ | llama-b11292-bin-win-cuda-13.4-x64.zip | cudart-llama-bin-win-cuda-13.4-x64.zip |
| NVIDIA GTX 10-series or older, or a driver below 580 | llama-b11292-bin-win-cuda-12.4-x64.zip | cudart-llama-bin-win-cuda-12.4-x64.zip |
| AMD Radeon (any recent card) | llama-b11292-bin-win-vulkan-x64.zip | GPU driver only |
| AMD Radeon, ROCm path | llama-b11292-bin-win-rocm-10.0-x64.zip | AMD ROCm / HIP SDK installed |
| Intel Arc or Intel iGPU | llama-b11292-bin-win-vulkan-x64.zip | GPU driver only |
| Intel GPU, oneAPI path | llama-b11292-bin-win-sycl-x64.zip | Bundled (sycl9.dll, MKL) |
| Intel CPU, GPU, or NPU via OpenVINO | llama-b11292-bin-win-openvino-2026.4-x64.zip | Bundled OpenVINO DLLs |
| No GPU, x64 | llama-b11292-bin-win-cpu-x64.zip | Nothing |
| Snapdragon X (Windows on ARM) | llama-b11292-bin-win-cpu-arm64.zip or win-opencl-adreno-arm64.zip | Nothing |
| Windows on ARM with an NVIDIA GPU | llama-b11292-bin-win-cuda-13.4-arm64.zip | cudart-llama-bin-win-cuda-13.4-arm64.zip |
A listing of the b11292 zips shows why the CPU row is simple. The x64 CPU zip ships separate CPU backend DLLs from ggml-cpu-sse42.dll through ggml-cpu-haswell.dll, ggml-cpu-zen4.dll, and ggml-cpu-sapphirerapids.dll (Source: llama.cpp releases).
Inference: llama.cpp picks the best CPU backend for your processor at startup, which is why no separate AVX2 or AVX-512 zip exists. For the rows marked "GPU driver only," the zip bundles no GPU runtime, so the Vulkan support in your normal graphics driver is what it relies on.
Older guides still send readers to asset names that no longer match. The Jan app's llama.cpp docs list names such as win-avx2-cuda-cu12.0-x64 (Jan docs). Those are Jan's own backend builds rather than upstream release assets. There is also no win-hip zip any more; the AMD build is now called win-rocm-10.0-x64 (Source: llama.cpp releases).
CUDA 12.4 or CUDA 13.4: which NVIDIA build?
The CUDA version you pick depends on your GPU generation and your driver, not on which CUDA Toolkit you have installed. The prebuilt zip plus its cudart zip carries the runtime it needs.
NVIDIA's CUDA 13 release notes say it "Removed support for Maxwell, Pascal, and Volta GPUs," which covers everything before the Turing generation. Turing is the RTX 20-series. The same notes say "Existing CUDA 13.x applications run on drivers >=580." By contrast, CUDA 12.4 pairs with Windows driver 551.61 or later (Source: NVIDIA CUDA release notes).
That gives a short rule you can follow:
- RTX 20, 30, 40, or 50-series with a current driver: use
cuda-13.4. - GTX 10-series, GTX 900-series, Titan V, or a driver older than 580: use
cuda-12.4. - Not sure: check the driver version in the NVIDIA app or run
nvidia-smi, then pick from the two rows above.
The cudart zips are named without a build number, cudart-llama-bin-win-cuda-13.4-x64.zip, and they contain exactly three files: cudart64_13.dll, cublas64_13.dll, and cublasLt64_13.dll. The 12.4 version holds the same three files with _12 names (Source: llama.cpp releases).
Inference: because the cudart zip has no build number, you only need to download it again when you switch between CUDA 12 and CUDA 13, not on every llama.cpp update.
Install llama.cpp on Windows: zip, winget, or install.ps1
There are three prebuilt routes, and they do not install the same thing. That is the part most guides skip.
| Route | Command | What you get | Best for |
|---|---|---|---|
| Release zip | Download from GitHub Releases | The exact backend you choose, plus llama.exe, llama-cli.exe, llama-server.exe | NVIDIA users and anyone who wants control |
| winget | winget install llama.cpp | The Vulkan x64 build, with llama-cli, llama-server, llama-bench on PATH | AMD and Intel users who want winget upgrade |
| llama.app script | irm https://llama.app/install.ps1 | iex | One llama.exe picked by probing CUDA, then ROCm, then Vulkan, then CPU | A single binary with llama serve and llama cli |
The winget manifest for package ggml.llamacpp points its x64 installer at llama-b11193-bin-win-vulkan-x64.zip. On ARM64 it points at the OpenCL Adreno zip. It also declares a dependency on the Microsoft Visual C++ 2015+ Redistributable (Source: winget manifest). So NVIDIA owners who use winget get the Vulkan backend, not CUDA. It runs, but if you bought an NVIDIA card for CUDA, take the release zip instead.
The llama.app PowerShell script installs llama.exe into %LOCALAPPDATA%\Microsoft\WindowsApps. Its CUDA probe stops with "NVIDIA GPU detected, but the CUDA Toolkit is not installed" when the toolkit is missing. The script then falls through to the ROCm, Vulkan, and CPU probes (Source: llama.app install.ps1).
The project's install docs also list conda-forge, which ships CUDA and Vulkan builds for Windows.
First run: download a model and start chatting
Open PowerShell in the folder where you extracted the zip. Then run the same command the project README uses:
.\llama-cli.exe -hf ggml-org/Qwen3.5-0.8B-GGUF
The -hf flag downloads a GGUF model from a Hugging Face repository. If you do not name a quant, it defaults to Q4_K_M, or the first file in the repo when that quant does not exist (Source: llama.cpp CLI reference). A 0.8B model is small enough to confirm the install works before you download something larger.
To serve the same model over an OpenAI-compatible API with the built-in web UI, one line is enough:
.\llama.exe serve -hf ggml-org/Qwen3.5-0.8B-GGUF
llama.exe serve and llama-server.exe start the same server, according to the llama dispatcher source; the zip ships both. The server listens on 127.0.0.1:8080 by default (server README). Server flags for context size, parallel slots, API keys, and Docker are a separate topic from getting the binary running.
When you move past the test model, the quant choice decides whether a model fits your VRAM. The trade-off between the two most common levels is covered in Q4_K_M vs Q5_K_M. If you want a background service and a curated model library instead of raw binaries, see Ollama vs llama.cpp.
For every server flag, the API endpoints, and the built-in web UI, see the llama-server guide.
Check that llama.cpp is using your GPU
A wrong zip does not fail loudly. It runs on the CPU at a fraction of the speed. Two checks take under a minute.
First, list the devices the build can see:
.\llama-cli.exe --list-devices
This prints the available devices and exits (Source: llama.cpp CLI reference). A working CUDA build shows a line such as CUDA0: NVIDIA GeForce .... An empty list, shown as Available devices: (none), means the backend did not load (issue #26929).
Second, load a model and read the startup log for the load_tensors lines. A user report on the issue tracker shows the pattern (issue #25488):
load_tensors: offloaded 22/41 layers to GPU
load_tensors: CUDA0 model buffer size = 10667.23 MiBYou do not need to pass -ngl for this. The GPU-layers option now defaults to auto, and --fit defaults to on, so llama.cpp sizes the offload to your free VRAM (Source: llama.cpp CLI reference).
Read the numbers as a diagnosis:
offloaded 41/41: the whole model is on the GPU.offloaded 22/41: the model is larger than your free VRAM. Use a smaller quant or model, or estimate the fit with the local LLM hardware calculator.offloaded 0/41or no GPU buffer line: you are on the CPU. Check the error table below.
Common llama.cpp Windows errors and fixes
These are the failures that recur on the llama.cpp issue tracker, each with the fix a primary source supports.
| Symptom | Cause | Fix |
|---|---|---|
cudart64_13.dll or cublas64_13.dll was not found | CUDA zip extracted without its runtime DLLs | Extract cudart-llama-bin-win-cuda-13.4-x64.zip into the same folder; use the 12.4 zip for _12 names (Source: llama.cpp releases) |
VCRUNTIME140.dll or MSVCP140.dll was not found | Visual C++ runtime missing | Install the Microsoft Visual C++ 2015+ Redistributable, the dependency winget declares (Source: winget manifest) |
CUDA zip runs, but no CUDA0 device appears | GPU older than Turing, or driver below 580, on the 13.4 build | Switch to the cuda-12.4 zip or update the driver (Source: NVIDIA CUDA release notes) |
ROCm zip shows Available devices: (none) | AMD ROCm / HIP SDK not installed | Install AMD's HIP SDK, or use the Vulkan zip, which users report working on the same cards (Source: llama.app install.ps1; #26929) |
llama is not recognized after winget install | winget exposes llama-cli, llama-server, and others, but not llama.exe | Run llama-cli or llama-server in a new terminal (Source: winget manifest) |
Defender quarantines a file as Trojan:Win32/Wacatac | A Defender machine-learning (!ml) detection on release binaries | Maintainers closed repeated reports as a false positive (#27716); verify with gh attestation verify <zip> -R ggml-org/llama.cpp before allowing it (Source: llama.cpp releases) |
llama update swapped a CUDA build or left a second copy | update runs install.ps1, which installs to WindowsApps (#24744) | Update zip installs by downloading the new zip; keep one install on PATH (Source: llama.app install.ps1) |
libomp140.aarch64.dll not found on Snapdragon | Self-built ARM64 binary without the OpenMP DLL (#28291) | Use the prebuilt win-cpu-arm64 zip, which bundles libomp.dll (Source: llama.cpp releases) |
Two AMD issues remain open. Issue #26964 reports ROCm builds falling back to CPU after the HIP-to-ROCm packaging change, while Vulkan works on the same RX 9070 XT. Inference: on AMD, Vulkan is the lower-risk default until those issues close.
AMD also hosts its own validated Windows builds. Its ROCm docs currently point to a b8407 zip built for ROCm 7.2.1 and gfx110X, gfx115X, and gfx120X GPUs (AMD ROCm docs). That build is thousands of releases behind upstream, so newer model architectures may not load.
When to build llama.cpp from source on Windows
Most readers never need to compile. Building makes sense when a prebuilt zip does not cover your hardware or you need a patch before it ships in a release.
The build guide lists the Windows requirements. You need Visual Studio 2022 with the C++ CMake tools, Git, and the Clang compiler components. Windows on ARM builds use the arm64-windows-llvm-release preset. The Vulkan backend needs the LunarG Vulkan SDK at build time (Source: llama.cpp build guide).
Build from source in these cases:
- Your AMD GPU target is missing from the ROCm zip. An issue reported the Windows HIP release omitted gfx1152 (#26127).
- You need a CUDA architecture flag the release does not set. The build guide documents
CMAKE_CUDA_ARCHITECTURESfor this (Source: llama.cpp build guide). - You want a fix merged hours ago. Releases are frequent, but a source build removes the wait.
The trade-off is maintenance. A zip is one download per update; a source build needs the toolchain kept current.
FAQ
How to install llama.cpp on Windows?
Download the release zip that matches your GPU from GitHub Releases: CUDA 13.4 for RTX 20-series and newer, Vulkan for AMD or Intel, CPU x64 with no GPU. NVIDIA users also extract the matching cudart-llama-bin-win-cuda zip into the same folder. Then run .\llama-cli.exe -hf ggml-org/Qwen3.5-0.8B-GGUF in PowerShell.
What is the best way to install llama.cpp on Windows?
For NVIDIA GPUs, the release zip is the best route, because winget installs the Vulkan build instead of CUDA. For AMD and Intel GPUs, winget install llama.cpp is simplest, and the package tracks new releases for winget upgrade. The llama.app script suits people who want one llama.exe with llama serve and llama cli subcommands.
Should I use CUDA or Vulkan for llama.cpp on Windows?
Use CUDA on NVIDIA cards, because CUDA is NVIDIA's own compute platform and llama.cpp ships dedicated CUDA 12.4 and 13.4 zips for it. Use Vulkan on AMD and Intel GPUs; it needs only the normal graphics driver. Vulkan also runs on NVIDIA, which is why a winget install still works on an RTX card, just not with the CUDA backend.
Does llama.cpp work on Windows on ARM?
Yes. Snapdragon X laptops can use the win-cpu-arm64 zip, or the win-opencl-adreno-arm64 zip to use the Adreno GPU, and winget installs the Adreno build on ARM64. Windows on ARM machines with an NVIDIA GPU have their own win-cuda-13.4-arm64 zip and a matching cudart zip.
Related coverage
References
- llama.app install.ps1 - https://llama.app/install.ps1
- llama.cpp build guide - https://github.com/ggml-org/llama.cpp/blob/master/docs/build.md
- llama.cpp CLI reference - https://github.com/ggml-org/llama.cpp/blob/master/tools/cli/README.md
- llama.cpp releases - https://github.com/ggml-org/llama.cpp/releases
- NVIDIA CUDA release notes - https://docs.nvidia.com/cuda/cuda-toolkit-release-notes/index.html
- winget manifest - https://github.com/microsoft/winget-pkgs/tree/master/manifests/g/ggml/llamacpp

