ThinkFacility Sign in

Error messages

Expected all tensors to be on the same device, but found at least two devices, cuda:0 and cpu!

The message

Expected all tensors to be on the same device, but found at least two devices, cuda:0 and cpu!
PyTorch 2.14.1 read October 1, 2026PyTorchComfyUIlocal models

What it means

One calculation got two inputs where one sat in GPU memory (cuda:0) and the other in system memory (cpu). PyTorch won't mix them, so it stopped. The code that forgot to move a tensor is usually a custom node, an extension or one model file.

What to do

In ComfyUI, start with --disable-all-custom-nodes; if the error goes, find the node and update it or report it to that node's repo. Update ComfyUI itself too, since a few core cases were bugs that got fixed.

A ComfyUI workflow gets as far as the KSampler or the VAE decode, or an A1111 generation starts, and the console prints:

RuntimeError: Expected all tensors to be on the same device, but found at least two devices, cuda:0 and cpu!

The two device names can come in either order, and older builds add a tail such as (when checking argument for argument mat1 in method wrapper_CUDA_addmm). As of October 1, 2026 there were 57 issues with this phrase in the ComfyUI repo, 39 in AUTOMATIC1111's Stable Diffusion WebUI and 15 in Forge.

What PyTorch is complaining about

The text comes from PyTorch, the library every one of these apps runs on. We read it at v2.14.1, in aten/src/ATen/TensorIterator.cpp. Before an operation runs, PyTorch checks that all its inputs (tensors, the arrays of numbers a model works on) live on one device. cuda:0 is your first Nvidia GPU and cpu is ordinary system RAM. If one input is in each, it stops with the line above.

Apart from single numbers, PyTorch doesn't move data between the two for you. Whatever code built that operation had to put both tensors on the GPU first, and one got left behind. Image tools shuffle models in and out of VRAM all the time (ComfyUI's own help text for --highvram says models are unloaded to CPU memory after use by default), so a node that grabs a model without moving it back can trip this.

The newer wording

PyTorch 2.8.0 changed one version of the check. When it can name the argument, it now prints this (from a ComfyUI training node's issue in September 2026):

Expected all tensors to be on the same device, but got weight is on cpu, different from other tensors on cuda:0 (when checking argument in method wrapper_CUDA___slow_conv2d_forward)

Here "weight" is a layer's weights, which were still in RAM. Before 2.8.0 the same check printed "found at least two devices ... (when checking argument for argument weight in method ...)".

Finding the culprit in ComfyUI

ComfyUI's maintainers have said more than once where these belong. ltdrdata wrote in #3179 (February 2025): "The device mismatch issue caused by a custom node cannot be resolved in the ComfyUI core. It needs to be addressed in the respective custom node repository." When an update broke PuLID with this error, his answer was that it's "a backward compatibility issue" for the PuLID node's repo to fix.

So test without custom nodes. Launch with --disable-all-custom-nodes, the method ComfyUI's troubleshooting docs describe. If a plain workflow works, turn nodes back on half at a time until it breaks, then update that node or open an issue on its repo.

Core bugs do happen. That same #3179 started as a workflow with several models and two ControlNets and no extra nodes, and comfyanonymous confirmed it in April 2024 and fixed it the next day. A low-VRAM LTX-Video decode that failed this way (#6144) got a fix commit in December 2024, and ltdrdata's first suggestion there was to turn off --bf16-vae. Updating ComfyUI is worth doing before you dig further.

In Stable Diffusion WebUI

Extensions are the usual suspect, so try with them disabled. One A1111 report (#16263, July 2024) came from picking a VAE downloaded from Hugging Face. A collaborator said that file was in the diffusers format, which the WebUI doesn't load, and told the user to set the VAE back to Automatic.

Other lines the same feature prints

Match yours against these if the one at the top of the page is not quite it. They come from the same code and mean related things.

  • RuntimeError: Expected all tensors to be on the same device, but found at least two devices, cuda:0 and cpu!
  • Expected all tensors to be on the same device, but found at least two devices, cpu and cuda:0!
  • RuntimeError: Expected all tensors to be on the same device, but found at least two devices, cpu and cuda:0! (when checking argument for argument mat1 in method wrapper_CUDA_addmm)
  • RuntimeError: Expected all tensors to be on the same device, but found at least two devices, cpu and cuda:0! (when checking argument for argument index in method wrapper_CUDA__index_select)
  • Expected all tensors to be on the same device, but got weight is on cpu, different from other tensors on cuda:0 (when checking argument in method wrapper_CUDA___slow_conv2d_forward)