Blog / Tutorials / Wan 2.2 Video Generation with ComfyUI: A GPU VPS Setup Guide
(last updated: )

Wan 2.2 Video Generation with ComfyUI: A GPU VPS Setup Guide

A step-by-step guide to running Wan 2.2 video generation in ComfyUI on a GPU VPS, from installing ComfyUI and downloading the model files to loading the native workflow templates and fixing common errors.

10 min read
Wan 2.2 Video Generation with ComfyUI: A GPU VPS Setup Guide

In short. Wan 2.2 is an open-weight video generation model from Wan AI, released under Apache 2.0, and ComfyUI supports it natively through built-in workflow templates. The 5B hybrid model, Wan2.2-TI2V-5B, runs on roughly 8 GB of VRAM under ComfyUI native offloading. The two 14B expert models need considerably more. A Contabo GPU VPS carries an NVIDIA RTX PRO 6000 with 96 GB of GDDR7 VRAM, so every Wan 2.2 variant fits on one card.

Video generation is the workload that finally makes a rented GPU worth the money. A still-image model will run on a laptop if you are patient. A 14B video model will not, and the gap between “technically possible” and “usable” is measured in hours per clip. This ComfyUI tutorial walks through the whole path on a remote GPU VPS: install, model download, the native Wan 2.2 ComfyUI workflow template, and the errors you will hit on the way.

What you need before you start

You need a GPU VPS with at least 8 GB of VRAM for the 5B model, an Ubuntu image with the NVIDIA driver and CUDA toolkit in place, root or sudo SSH access, and enough free disk for the weights.

  • A GPU VPS, a virtual server with a dedicated NVIDIA GPU passed through to it. The Contabo GPU VPS ships Ubuntu 24.04 LTS with the CUDA drivers, the CUDA toolkit, and nvidia-container-toolkit already integrated, which removes the driver install that usually eats the first hour.
  • SSH access as root or a sudo user. Everything below runs over SSH.
  • Free disk space. The 5B diffusion model is about 10 GB. The 14B text-to-video pair in fp8 runs to roughly 29 GB across both experts, and the fp16 image-to-video pair to about 57 GB. The 900 GB of NVMe on a GPU VPS absorbs all of them with room left over.
  • An SSH client locally. You will tunnel the ComfyUI interface back to your own browser rather than exposing it.

Confirm the card is visible before you install anything:

nvidia-smi

If that prints a table with the GPU name and a driver version, you are ready. If it prints nothing, stop and fix the driver first.

Installing ComfyUI on a Contabo GPU VPS

Installing ComfyUI takes four steps, then two more to run it: clone the repository, create a Python virtual environment, install a CUDA-matched PyTorch build, and install the remaining dependencies. The rest of this ComfyUI tutorial assumes Ubuntu 24.04 and a working NVIDIA driver.

  1. Install the system packages ComfyUI needs.
sudo apt update
sudo apt install -y git python3-venv python3-pip
  1. Clone the repository and create an isolated environment.
git clone https://github.com/comfyanonymous/ComfyUI.git
cd ComfyUI
python3 -m venv venv
source venv/bin/activate
  1. Install PyTorch against the right CUDA version, then the ComfyUI requirements. This is the step that goes wrong most often. ComfyUI requires a cu130 or newer PyTorch build on any NVIDIA 20-series card or later, which includes the RTX PRO 6000. A default pip install torch may pull an older wheel that installs fine and then reports no usable GPU.
pip install torch torchvision torchaudio --extra-index-url https://download.pytorch.org/whl/cu130
pip install -r requirements.txt
  1. Verify that PyTorch can see the card.
python -c "import torch; print(torch.cuda.is_available(), torch.cuda.get_device_name(0))"

The output should read True followed by the GPU name. Anything else means the wheel and the driver disagree, and no amount of workflow tuning will fix it.

  1. Start the server, bound to localhost.
python main.py --listen 127.0.0.1 --port 8188
  1. From your own machine, open an SSH tunnel and then load the interface at http://127.0.0.1:8188.
ssh -L 8188:127.0.0.1:8188 root@YOUR_SERVER_IP

Binding to 127.0.0.1 and tunneling keeps ComfyUI off the public internet. ComfyUI has no authentication of its own, so an instance bound to 0.0.0.0 is an open remote code execution endpoint for anyone who finds the port.

Downloading the Wan 2.2 model files

A Wan 2.2 workflow in ComfyUI needs three file types: a diffusion model, a VAE, and a text encoder. All of them live in the Comfy-Org/Wan_2.2_ComfyUI_Repackaged repository on Hugging Face, already converted to the layout ComfyUI expects.

For the 5B hybrid model, which handles both text-to-video and image-to-video:

cd ~/ComfyUI
wget -P models/diffusion_models https://huggingface.co/Comfy-Org/Wan_2.2_ComfyUI_Repackaged/resolve/main/split_files/diffusion_models/wan2.2_ti2v_5B_fp16.safetensors
wget -P models/vae https://huggingface.co/Comfy-Org/Wan_2.2_ComfyUI_Repackaged/resolve/main/split_files/vae/wan2.2_vae.safetensors
wget -P models/text_encoders https://huggingface.co/Comfy-Org/Wan_2.1_ComfyUI_repackaged/resolve/main/split_files/text_encoders/umt5_xxl_fp8_e4m3fn_scaled.safetensors

The text encoder comes from the Wan 2.1 repackaged repository and is shared across both generations. That is expected, not a broken link.

When the downloads finish, the tree should look like this:

ComfyUI/
  models/
    diffusion_models/
      wan2.2_ti2v_5B_fp16.safetensors
    text_encoders/
      umt5_xxl_fp8_e4m3fn_scaled.safetensors
    vae/
      wan2.2_vae.safetensors

The 14B text-to-video workflow needs a different set: wan2.2_t2v_high_noise_14B_fp8_scaled.safetensors, wan2.2_t2v_low_noise_14B_fp8_scaled.safetensors, and wan_2.1_vae.safetensors. The image-to-video workflow swaps in the i2v equivalents. Restart ComfyUI after adding files so the loader nodes pick them up.

Loading the Wan 2.2 workflow in ComfyUI

ComfyUI ships every Wan 2.2 ComfyUI workflow as a built-in template, so you do not build the graph by hand. Open the workflow menu, browse the template library, and search for “Wan2.2 5B” to load the hybrid text-to-video and image-to-video graph.

Once the template opens, check four node bindings before you run anything:

  1. The Load Diffusion Model node points at wan2.2_ti2v_5B_fp16.safetensors.
  2. The Load CLIP node points at umt5_xxl_fp8_e4m3fn_scaled.safetensors.
  3. The Load VAE node points at wan2.2_vae.safetensors.
  4. The CLIP Text Encode nodes hold your positive and negative prompts.

Resolution and clip length live in the Wan22ImageToVideoLatent node in the 5B template, and in EmptyHunyuanLatentVideo in the 14B templates. The length field is frame count, not seconds. For the 5B model, 720p means 1280x704 or 704x1280, not 1280x720, because of how the Wan 2.2 VAE compresses. To do image-to-video instead of text-to-video with the 5B model, select the Load Image node and press Ctrl+B to enable it, then upload your starting frame.

Press Ctrl+Enter to queue the job. If a node loads red with a “missing node type” error, your ComfyUI checkout predates the Wan 2.2 nodes. A git pull followed by reinstalling requirements resolves it.

ComfyUI ships two other Wan 2.2 templates. “Wan2.2 14B I2V” animates a single input image, and “Wan2.2 14B FLF2V” interpolates between a defined first and last frame. The FLF2V template ships at a deliberately small default resolution so low-VRAM machines can run it. On 96 GB of VRAM you can raise it toward 720p.

Choosing between the 5B and 14B models

The 5B model is the one to start with, and the 14B pair is what you graduate to once the pipeline works end to end. Wan 2.2 uses a mixture-of-experts design at 14B, splitting generation between a high-noise expert and a low-noise expert across denoising timesteps, which is why those workflows load two diffusion models instead of one.

ModelParametersTaskVRAM guidance
Wan2.2-TI2V-5B5BText-to-video and image-to-video, 720p at 24 fpsAround 8 GB with ComfyUI native offloading
Wan2.2-T2V-A14B14B (MoE)Text-to-video at 480p and 720p, cinematic aesthetic control80 GB or more in the Wan reference implementation, lower with the ComfyUI fp8 weights
Wan2.2-I2V-A14B14B (MoE)Image-to-video at 480p and 720p, motion from a still frame80 GB or more in the Wan reference implementation, lower with the ComfyUI fp8 weights

The practical read: 8 GB of VRAM gets you working output, while 96 GB removes the ceiling on resolution and clip length entirely. All Wan 2.2 weights are Apache 2.0 licensed, so commercial use is permitted as long as you keep the copyright notice and license text.

Common errors and how to fix them

torch.cuda.is_available() returns False, or ComfyUI reports “Torch not compiled with CUDA enabled,” while nvidia-smi works. The PyTorch wheel does not match the card. Run pip uninstall torch and reinstall from the cu130 index. On recent NVIDIA cards this is the single most common failure.

Red nodes and “missing node type” when the template loads. Your ComfyUI version is older than the Wan 2.2 node set. Run git pull inside the ComfyUI directory, reinstall requirements inside the virtual environment, and restart the server.

CUDA out of memory partway through sampling. Reduce the frame count first, then the resolution. Frame count drives peak VRAM harder than width and height do. If you are on the 14B workflow and still hitting the wall, drop back to the 5B template.

The browser cannot reach port 8188. Check that the SSH tunnel is still open and that the server is running. Do not solve this by rebinding ComfyUI to 0.0.0.0 and opening the firewall.

Why run Wan 2.2 on a Contabo GPU VPS

A Contabo GPU VPS gives you a full NVIDIA RTX PRO 6000 through 1:1 PCI passthrough, with 96 GB of GDDR7 VRAM, 18 vCPUs, 96 GB of RAM, and 900 GB of NVMe storage on a flat monthly bill. There is no vGPU slicing, so the card you benchmark is the card you get for the whole month.

The 96 GB figure is what makes this card interesting for Wan 2.2 specifically. The Wan reference implementation asks for 80 GB of VRAM to run a 14B model on a single GPU, which rules out every consumer card and most of the mid-range rental market. A 96 GB card clears that bar without multi-GPU sharding or a single line of distributed inference code.

That billing model matters more for video than for most workloads. Video generation is bursty: you iterate on prompts for an afternoon, leave the box idle overnight, then render a batch. Per-second hyperscaler billing punishes the idle time, and per-second reservation queues punish the burst. A flat rate does neither.

Contabo supplies the GPU VPS and the Ubuntu image with CUDA already in place. ComfyUI, the Wan 2.2 weights, and everything you generate with them are yours to install and maintain. Nobody at Contabo touches the instance, and nothing you render passes through a managed service. For a model whose license explicitly permits commercial output, that separation is the point.

Data centers are available in Europe and US Central, which keeps generation traffic inside a jurisdiction you choose.

FAQ: Wan 2.2 video generation on a GPU VPS

How much VRAM does Wan 2.2 need?

The 5B hybrid model, Wan2.2-TI2V-5B, fits in roughly 8 GB of VRAM under ComfyUI native offloading. The 14B models are heavier: the Wan reference implementation asks for at least 80 GB of VRAM for single-GPU A14B inference, though the fp8 weights ComfyUI ships lower that figure. A 96 GB card clears the official requirement outright.

What’s the easiest way to run a ComfyUI workflow in the cloud?

Rent a GPU VPS with a CUDA-ready Ubuntu image, install ComfyUI into a Python virtual environment, bind the server to localhost, and reach it through an SSH tunnel from your own browser. That gives you the full ComfyUI workflow editor with no authentication layer exposed to the internet, and the models stay on your disk rather than a third-party queue.

Is Wan 2.2 better than Wan 2.1?

Wan 2.2 adds a mixture-of-experts architecture that splits denoising between a high-noise and a low-noise expert, and Wan AI reports training it on 65.6% more images and 83.2% more videos than Wan 2.1. It also adds the 5B hybrid variant, which handles text-to-video and image-to-video in one model. Wan 2.1 remains usable, and the two share a text encoder.

Can I use Wan 2.2 output commercially?

Yes. The Wan 2.2 model series is released under the Apache 2.0 license, which permits commercial use, modification, and redistribution provided you retain the original copyright notice and license text. That covers the weights themselves. Check your own footage, reference images, and any LoRA you add to the workflow separately, since those carry their own terms.

Do I need ComfyUI Manager to run Wan 2.2?

No. The Wan 2.2 workflows ship as native ComfyUI templates and use core nodes only, so a current checkout runs them without extra packages. ComfyUI Manager now lives in the ComfyUI repository: install <code>manager_requirements.txt</code>, then start the server with the <code>–enable-manager</code> flag. You need it once you start adding community custom nodes, such as the GGUF loader or the WanVideoWrapper set.