gpu-vps-ai-training_hero-background-image_alt
icon-gpu_f_white_alt
GPU VPS

Dedicated GPU VPS for AI Training & Inference

Run your own LLMs and AI models on a dedicated NVIDIA RTX PRO 6000 with 96 GB of VRAM. Fine-tune a model, serve inference, or build a RAG pipeline minutes after checkout. CUDA is pre-installed and the whole GPU is yours, with no sharing and no GPU slicing.


Dedicated GPU resources. Monthly billing available. No hidden costs. NVMe storage.

Object Storage
GPU VPS
€999.00
799
20 / month
 

Dedicated NVIDIA RTX PRO 6000

CUDA-compatible・NVMe storage

Unlimited traffic*


DEDICATED NVIDIA RTX 6000 PRO
FOR EVERY AI WORKLOAD

Performance VPS meets NVIDIA GPU power with the latest Blackwell architecture

and native FP4 for quantized models. Go beyond consumer cards

with enough resources to load serious LLMs.

gpu animation
96 GB
GDDR7 VRAM
24,064
CUDA CORES
120
FP32 TFLOPS
1,597
GB/s Bandwidth

Get started in minutes

No manual driver installs. No long provisioning queue. Just a working CUDA environment, fast.

Step 1: Choose your Location

Pick your GPU VPS region (Europe or US Central).

Step 2: Finalize your order

Your instance deploys automatically on Ubuntu 24.04 with CUDA drivers, the CUDA toolkit, and nvidia-container-toolkit pre-integrated.

Step 3: Connect and start working

SSH in and go. With the basic setup done for you, you can launch any CUDA-compatible framework the moment you connect.

Why Run AI Training and Inference Workloads on a Contabo GPU VPS?

background-image_gpu-vps-why-run-ai-training_alt
icon-dedicated-servers_f_alt
100% Dedicated GPU Resources

The full RTX PRO 6000 is mapped to a Performance VPS with no vGPU slicing. All 96 GB stays yours, so throughput holds from the first training step to the last.

icon-platform_f_alt
CUDA-Compatible Infrastructure

Every GPU VPS comes with Ubuntu 24.04 with CUDA already installed. Pull a PyTorch or vLLM container and get straight to work.

icon-enlarge_f_alt
Built for Large Models

Blackwell architecture with native FP4 and 96 GB of VRAM runs models a consumer card cannot hold. A 70B-class model serves at FP8 on a single card.

icon-wallet_f_alt
Transparent Pricing

One flat monthly rate (annual billing available) with unlimited traffic (fair usage policy applies). No per-second or per-token metering.

icon-access_f_alt
Your Infrastructure on Your Terms

Root access and standard Linux tooling mean total control. Keep your weights, prompts and user data in a European or US data center you choose.

icon-support_f_alt
24/7 Support

Award-winning support from both human specialists and high-quality AI-assisted responses. Quick answers when a training run or an endpoint needs them.

How the RTX 6000 Pro Measures Up

FP32 measures single-precision compute, one baseline for a GPU's raw throughput. At 120 TFLOPS the RTX PRO 6000 Server Edition leads every data center and professional GPU compared here, and pairs that with 96 GB of VRAM and native FP4 for AI workloads.
RTX PRO 6000
120
Blackwell Server Edition · 96 GB GDDR7 ECC · 24,064 CUDA cores
L40S
91.6
Ada Lovelace · 48 GB GDDR6 · 18,176 CUDA cores
H100
67
Hopper · 80 GB HBM3 · 16,896 CUDA cores · built for multi-node training
H200
67
Hopper · 141 GB HBM3e · 16,896 CUDA cores · built for multi-node training
A100
19.5
Ampere · 80 GB HBM2e · 6,912 CUDA cores · previous generation
FP32 compute (TFLOPS) - higher is better

GPU VPS Use Cases for AI Teams

Four AI workloads that fit on a single 96 GB card.
icon-ai_f_alt
LLM Inference with Your Own Model

Serve your own weights and keep every prompt on infrastructure you control. A 70B-class model fits at FP8, with room left for context.

icon-configuration_f_alt
Fine-Tuning and LoRA Training

LoRA y Fine-Tuning de parámetros completos en una sola tarjeta. Los 96 GB de VRAM le permiten cargar modelos base que una GPU de consumo de 24 GB no puede cargar.

icon-merge_f_alt
RAG Pipelines at Scale

Run the embedder, vector store and generation model on one GPU. No network switching between them, and your data stays in your environment.

icon-team_f_alt
Dev & Test Environments for AI Teams

One always-on GPU box for the whole team, pre-configured with CUDA. No cold starts and no re-provisioning between sessions.

Customize Your GPU VPS Your Way with Apps and Deployment Options

GPU VPS comes with Ubuntu 24.04 and CUDA pre-installed. You can easily add a range of apps and panels using the Customer Control Panel.


vps-page-tab-1-click-openclaw
OpenClaw
vps-page-tab-1-click-openclaw
n8n
logo-small-hermes_alt
Hermes Agent
logo_small_ollama
Ollama
logo-small-dokploy_alt
Dokploy
logo-small-paperclip_alt
Paperclip
logo-small-coolify_alt
Coolify
Logo cPanel small

cPanel

vps-page-tab-panel-plesk
Plesk
vps-page-tab-panel-docker
Docker
vps-page-tab-panel-webmin
Webmin
vps-page-tab-panel-proxmox
Proxmox
icon-custom-images_f_alt

Deploy Your Perfect Image

With Custom Images, you can instantly deploy your .iso or .qcow2 image via web UI or API.

image-hugging-face_alt

Looking for inspiration? Browse AI
models and apps on Hugging Face.

FAQs

Frequently Asked Questions about our GPU VPS
What is a GPU VPS?

A GPU VPS is a Virtual Private Server with a dedicated NVIDIA GPU attached via PCI passthrough. Unlike shared GPU instances, the full GPU, including VRAM, CUDA cores, and memory bandwidth, is reserved for your workload alone.

Which GPU does Contabo GPU VPS use?

Every GPU VPS runs the NVIDIA RTX PRO 6000 Blackwell Server Edition, with 96 GB of GDDR7 ECC VRAM, 24,064 CUDA cores, 120 TFLOPS of FP32 compute, and 1,597 GB/s of memory bandwidth.

Is the GPU dedicated or shared on Contabo GPU VPS?

The GPU is dedicated. The full RTX PRO 6000 is allocated to your VPS through 1:1 PCI passthrough, with no vGPU slicing.


Is this GPU VPS good for training, or only for inference?

Both. A dedicated RTX PRO 6000 with 96 GB of VRAM and 5th-generation Tensor Cores handles training, fine-tuning and inference. Fine-tuning and adapting existing models is its strongest use, which covers most real-world training work.


Can I run LLM inference and fine-tuning on the same GPU VPS?

Yes. The usual workflow is to fine-tune your model on the GPU VPS, then serve it for inference on the same dedicated RTX PRO 6000. Both are single-GPU workloads, and 96 GB of VRAM gives room to move from one stage to the next.


Can I run multi-GPU or multi-node AI training on a Contabo GPU VPS?

A single GPU VPS is built for fine-tuning and adapting existing models on one dedicated RTX PRO 6000 with 96 GB of VRAM. For large-scale training across multiple GPUs or nodes, our sister company vshosting offers dedicated multi-GPU servers with NVLink, H200 and B300 configurations.


Can a 70B parameter model actually run on this GPU VPS?

Yes, at FP8 precision. A 70B-class model's weights occupy roughly 70 GB at FP8, one byte per parameter, leaving headroom within 96 GB of VRAM for context and inference overhead.


Which AI models are confirmed to run on GPU VPS?

Stable Diffusion, Qwen3-14B, Nemotron-Nano-3-30B, Devstral-2-123B, Qwen-Image-2512, and Coqui TTS are all confirmed. Any other CUDA-compatible model that fits within 96 GB of VRAM will run.


Is CUDA pre-installed?

Yes. Every GPU VPS ships on Ubuntu 24.04 with CUDA and nvidia-container-toolkit already installed.


What operating systems are supported?

Ubuntu 24.04 LTS with CUDA pre-installed is the only supported image for GPU VPS in the order process. Windows can be installed later using the Customer Control Panel if desired - in this case, you will need to configure CUDA yourself.


Where are GPU VPS data centers located?

Europe and US Central are the only available regions at launch. Starting in two locations is a deliberate first step. GPU capacity is harder to plan than standard VPS, so we are keeping the launch focused to protect availability and performance, with more regions planned.


The price displayed is the effective monthly rate for a 24-month subscription including applicable taxes.

Minimum initial contract length: 1 month | Minimum following contract length: Equal to the initial contract length | Minimum cancellation notice: None (pre-paid) or 4 weeks (post-paid).

Incoming Traffic: Unlimited and unmetered. No extra charges apply. | Outgoing Traffic: Unlimited - fair usage policy applies based on average server workloads. To maintain fair network performance for all customers, Contabo reserves the right to throttle servers with exceptionally high or disruptive usage patterns.

Applications and Add-Ons: Contabo provides full technical assistance for the Contabo VPS itself - including the network, hardware, and initial setup of any 1-click Add-Ons. Any additional open source software included in the applications is owned by the respective provider - maintenance, upgrading, and troubleshooting are within the end-users responsibility.

Leave your details here and we'll notify you when GPU VPS becomes available.

All set! We'll notify you as soon as GPU VPS is available again.
Something went wrong. Please try again.