In short. KoboldCpp is a free, single-file program that runs GGUF language models on your own server and serves them through a web interface and an API. On a VPS you download one binary, add a model file, and launch it inside tmux. A CPU-only server with 11 GiB of RAM runs a 0.5 billion-parameter test model, and 7-8 billion parameters is the estimated limit.
What is KoboldCpp used for?
KoboldCpp is used to run open-weight language models on hardware you control, mostly for story writing, roleplay, and chat, and as a backend for other AI tools. It builds on llama.cpp, the C++ inference engine behind many local LLM tools, and is inspired by the original KoboldAI project. A developer known as Concedo (LostRuins on GitHub) maintains it under the AGPLv3 license.
The program ships as a single self-contained file, loads GGUF model files, and bundles KoboldAI Lite, a browser interface with story, chat, instruct, and adventure modes. It also offers an OpenAI-compatible API, so tools such as SillyTavern or your own scripts can use it as a backend. Image, speech, and music features exist as well. They are optional and not covered here.
Running it on a VPS gives you an always-on private endpoint. KoboldCpp runs the model on your own server, and its AI Horde worker, which shares your compute with other users, is opt-in. A small team can also share one instance, because KoboldCpp queues incoming requests.
What do you need before you start?
You need a Linux VPS, an SSH client on your own computer, and about 1 GB of free disk space for the binary and the test model.
- Tested environment: Ubuntu 24.04 with 6 vCPUs and 11 GiB of RAM, running KoboldCpp 1.122.1.
- Access: a root login. If you use a regular user, put
sudoin front of theaptcommand in Step 3. - SSH client: the steps marked “on your PC” use PowerShell on Windows. macOS and Linux terminals work the same way.
RAM decides which models fit, because the model file must fit in memory with room left for context. This guide starts with a 0.5 billion-parameter model, which works as a quick smoke test. As an estimate, a 7-8 billion-parameter model at 4-bit quantization is the realistic limit on 11 GiB. CPU generation is slow, so treat this setup as a way to try the stack, not as a chat service for many users.
How do you deploy KoboldCpp on a VPS?
You deploy KoboldCpp in eight steps: connect, check resources, install tools, download the binary and a model, start it in tmux, verify, and open the web interface through an SSH tunnel.
Step 1: Connect (on your PC). Run ssh root@YOUR_SERVER_IP and type yes at the fingerprint prompt. If SSH warns that the remote host identification has changed, which happens after you reinstall the VPS, remove the old key first and reconnect:
ssh-keygen -R YOUR_SERVER_IP
ssh root@YOUR_SERVER_IP
Step 2: Check resources (on the VPS). This shows memory, disk space, and CPU count:
free -h && df -h / && nproc
Step 3: Install tools and create folders.
apt update && apt install -y curl tmux
mkdir -p ~/koboldcpp/models && cd ~/koboldcpp
Step 4: Download KoboldCpp. The KoboldCpp 1.122.1 release notes list three Linux builds. This guide uses the CPU build, which is about 130 MB.
| File | Use it for |
|---|---|
koboldcpp-linux-x64 | Servers with an NVIDIA GPU (CUDA build) |
koboldcpp-linux-x64-nocuda | Servers without an NVIDIA GPU, smaller download |
koboldcpp-linux-x64-oldpc | Older CPUs or older NVIDIA GPUs (CUDA 11 and AVX1) |
curl -fLo koboldcpp https://github.com/LostRuins/koboldcpp/releases/latest/download/koboldcpp-linux-x64-nocuda
chmod +x koboldcpp
Newer releases may change flags. To pin the tested version, replace latest/download in the address with download/v1.122.1.
Step 5: Download the test model. The model is Qwen2.5-0.5B-Instruct in Q4_K_M quantization. On Hugging Face, a file’s page address contains /blob/, which is a preview page. Replace /blob/ with /resolve/ to get the direct download link:
curl -fLo models/test.gguf "https://huggingface.co/Qwen/Qwen2.5-0.5B-Instruct-GGUF/resolve/main/qwen2.5-0.5b-instruct-q4_k_m.gguf"
ls -lh models
The listing should show test.gguf at about 469M. A file of a few kilobytes means you downloaded a /blob/ page.
Step 6: Start KoboldCpp in tmux. The tmux session keeps KoboldCpp running after you close SSH.
tmux new -s kobold
./koboldcpp --model ./models/test.gguf --usecpu --host 127.0.0.1 --port 5001 --contextsize 2048
The flags do five jobs: --model names the GGUF file, --usecpu keeps the work on the processor, --host sets the listening address, --port sets the port, and --contextsize 2048 limits memory for context, which defaults to 16384. Always set --host. In version 1.122.1, a launch without it also accepted connections on the server’s network address, even though the log prints localhost. The log ends like this when the model is ready:
Load Text Model OK: True
Please connect to custom endpoint at http://127.0.0.1:5001
Step 7: Detach and verify. Press Ctrl+B, release both keys, then press D. KoboldCpp keeps running in the background. Then check that it answers:
curl http://localhost:5001/api/extra/version
The reply starts with {"result": "KoboldCpp", "version": "1.122.1" and continues with more fields.
Step 8: Open the web interface (on your PC). In a new PowerShell window, open an SSH tunnel, which forwards port 5001 on your computer to the private port on the server:
ssh -L 5001:localhost:5001 root@YOUR_SERVER_IP
Keep that window open, then browse to http://localhost:5001.
How do you use and stop the running server?
You use it through three addresses on your tunnel, and you manage it with tmux. http://localhost:5001/ opens KoboldAI Lite, http://localhost:5001/lcpp/ opens the bundled llama.cpp interface, and http://localhost:5001/api shows the API documentation. To test generation, choose Instruct in the Format dropdown under Basic Settings and send a short prompt. The log also lists an OpenAI-compatible API at /v1/, which this guide does not test.
| Task | Command |
|---|---|
| Detach and leave it running | Ctrl+B, release, then D |
| Reattach | tmux attach -t kobold |
| Stop KoboldCpp | Reattach, then press Ctrl+C |
| Restart | Press the up arrow in the same session and run the launch command again |
| Remove the session | tmux kill-session -t kobold |
Does KoboldCpp need a CPU or a GPU?
KoboldCpp runs on a CPU alone, but a GPU makes generation faster once models get large. On a CPU, the model weights sit in system RAM and the processor reads them for every token. With GPU offloading, layers move into the card’s VRAM, where parallel hardware processes them much faster.
Three flags control this. --usecpu keeps everything on the processor, --usecuda or --usevulkan turns on GPU acceleration, and --gpulayers sets how many layers move to the card. The --help text in 1.122.1 lists -1 as the default for --gpulayers, which lets KoboldCpp choose automatically. Start on a CPU if your model is small and quantized and you can live with slow replies. Move to a GPU when you want answers that feel interactive or need a larger model.
Size the server by the model you plan to run, because the model file and its context must fit in memory. The KoboldCpp wiki gives a rough guide for 4-bit quantized models at a 2048-token context: at least 8 GB of RAM for a 7-billion-parameter model, 16 GB for 13 billion, 32 GB for 30 billion, and 64 GB for 65 billion. The default context in version 1.122.1 is 16384, so leave headroom or lower --contextsize. For a CPU-only setup, our Cloud VPS products are a good fit – choose depending on RAM. Choose Performance VPS if you want NVMe storage for faster model loading, or Max Performance VPS if you want dedicated resources for steadier generation speed. More cores also help, because the wiki advises setting the thread count close to the number of physical cores.
For larger models or faster replies, GPU VPS pairs a dedicated NVIDIA RTX PRO 6000 (96 GB of VRAM, CUDA-compatible) with 18 vCPUs and 96 GB of RAM, so the weights sit in VRAM instead of system RAM. As an estimate, a 70-billion-parameter model needs about 40 GB for its weights at 4-bit quantization, plus memory for context, so it fits on that one card.
What if something goes wrong?
Most problems come from SSH, model downloads, or memory. This table covers the ones that came up during testing.
| Symptom | Likely cause | Fix |
|---|---|---|
| SSH warns that the remote host identification has changed | You reinstalled the VPS, so the server key changed | Run ssh-keygen -R YOUR_SERVER_IP, then connect again |
| The model file is only a few kilobytes | You used a /blob/ link | Replace /blob/ with /resolve/ and download again |
| The tunnel fails with “Address already in use” | Port 5001 is busy on your computer | Use ssh -L 5002:localhost:5001 root@YOUR_SERVER_IP and browse to http://localhost:5002 |
Killed appears and KoboldCpp exits | The server ran out of memory | Pick a smaller model or quantization, or lower --contextsize |
| The web interface does not load in your browser | The tunnel window is closed, or KoboldCpp is not running | Keep the SSH window open and repeat the check from Step 7 |
What next?
Try a larger model, up to the limit of your RAM, and compare its replies with the test model. Speed drops as models grow, so watch how long a reply takes before you commit to one. If you open the port instead of using a tunnel, set --password and restrict the port to your own IP address in the Contabo Firewall. Running KoboldCpp as a systemd service is also an option, but this guide does not cover it.
FAQ: self-hosting KoboldCpp
Does KoboldCpp need a GPU to run?
No. KoboldCpp runs on a CPU alone when you launch it with --usecpu, which makes a RAM-rich VPS a workable host for small models. A Cloud VPS with CPU can be a comfortable fit for normal models and small requirements. A GPU, used through --usecuda or --usevulkan with --gpulayers, is the route to larger models and faster replies.
Is KoboldCpp free?
Yes. KoboldCpp is free, open-source software released under the AGPLv3 license, and you can download the prebuilt binary or compile it from source at no cost. Your only expenses are the server you run it on and the time to set it up. Model files are separate downloads, and each model carries its own license, so check the terms before using one commercially.
What model formats does KoboldCpp support?
For text models, KoboldCpp loads GGUF files and stays compatible with the older GGML format. Other formats such as safetensors and PyTorch .bin files are not supported natively and must be converted to GGUF first. You can find ready-made GGUF files on Hugging Face, usually in several quantization levels, from small and fast variants to larger, higher-quality ones.
How do I keep a KoboldCpp server private?
Set --host 127.0.0.1 and reach the server through an SSH tunnel. In version 1.122.1, a launch without --host also accepted connections on the server’s network address, even though the log printed localhost. If you open the port anyway, set --password, which protects text endpoints but not image endpoints, and restrict the port to your own IP address in the Contabo Firewall.
