CUDA_HOME environment variable is not set. Please set it to your CUDA install root.
The message
CUDA_HOME environment variable is not set. Please set it to your CUDA install root.What it means
Something tried to compile CUDA code against PyTorch, and PyTorch couldn't find the CUDA toolkit (the kit with Nvidia's nvcc compiler) anywhere it looks. A working GPU and a CUDA build of torch aren't enough for this; the compiler has to be installed too.
What to do
Install the CUDA toolkit whose major version matches torch.version.cuda and point CUDA_HOME at it (or put its nvcc on your PATH). For flash-attn on Linux, a prebuilt wheel from its releases page avoids compiling at all.
You run pip install flash-attn, or install something that builds its own GPU code, and pip fails with a long log and metadata-generation-failed. Near the bottom:
OSError: CUDA_HOME environment variable is not set. Please set it to your CUDA install root.
It confuses people because their GPU works fine. As of October 1, 2026 the phrase was in 469 GitHub issues, 13 of them in the flash-attention repo.
Where PyTorch looks for CUDA
The message is PyTorch's. We read torch/utils/cpp_extension.py at v2.14.1, the code packages use to compile their own CUDA extensions against torch. It searches in this order:
- the
CUDA_HOMEorCUDA_PATHenvironment variable; nvcc(Nvidia's CUDA compiler) on your PATH, taking the folder above itsbin;/usr/local/cudaon Linux;- the first
C:/Program Files/NVIDIA GPU Computing Toolkit/CUDA/v*.*folder on Windows.
If none exists, nothing happens yet. The code's own comment calls it "a lazy way of raising an error" only when a CUDA path is needed, which is the moment a build asks for CUDA's include or library folders.
Running a model and compiling for one are different jobs. The torch you installed with pip runs on your GPU without the toolkit, but it doesn't bring nvcc. flash-attn's setup script says the same about the official PyTorch Docker images: "only images whose names contain 'devel' will provide nvcc."
Why flash-attn hits it even with prebuilt wheels
We read flash-attn's setup.py at v2.8.3.post1. It tries to download a matching prebuilt wheel from its GitHub releases, chosen by the CUDA version your torch was built with. A missing toolkit only produces a warning there, written "because user could be downloading prebuilt wheels." But the script still describes a CUDAExtension, and creating that asks PyTorch for CUDA paths. Tri Dao, the maintainer, put it this way in #509 (September 2023): "Constructing the CUDAExtension with pytorch already requires CUDA_HOME."
His workaround in that thread was FLASH_ATTENTION_SKIP_CUDA_BUILD=TRUE pip install flash-attn --no-build-isolation, for setups where a prebuilt wheel exists. Prebuilt wheels are Linux only. The README says Windows "might work" from v2.3.2 but still needs more testing, and Windows reports like #982 run long.
The fix: install the toolkit and point at it
- Check which CUDA your torch was built for:
python -c "import torch; print(torch.version.cuda)". - Install the CUDA toolkit with the same major version. DeepSpeed's install docs say the major versions must match, and that a minor mismatch "may still" cause errors.
- Set
CUDA_HOMEto its root, such as/usr/local/cuda-12.8, or put itsbinfolder (withnvcc) on your PATH. - Check with
nvcc -V, then install again with--no-build-isolation, as flash-attn's README does.
Install ninja too. Without it, the README warns, compiling can take two hours.
DeepSpeed and other packages
DeepSpeed's docs say its ops are built "just-in-time" by default, using the same PyTorch loader. So pip install deepspeed can finish and the error can still turn up later, when an op is compiled. The fix is the same toolkit. The phrase also appears in 6 vLLM issues and a couple in AUTOMATIC1111's and ComfyUI's repos.
Other lines the same feature prints
Match yours against these if the one at the top of the page is not quite it. They come from the same code and mean related things.
OSError: CUDA_HOME environment variable is not set. Please set it to your CUDA install root.CUDA_HOME environment variable is not set