SGLang is an open-source serving framework that runs large language models quickly on your own GPU. It handles batching, caching, and structured output, then exposes an OpenAI-compatible API. Teams self-host SGLang to keep prompts and data on infrastructure they control, to avoid per-token fees, and to tune performance for their own workload.
What SGLang Does
A language model is a large file of weights. Loading it in a basic script is enough to generate text, but that approach handles one request, or a small fixed batch, at a time. When several users send prompts at once, requests queue up, GPU memory goes unused, and latency climbs.
SGLang closes that gap with a dedicated serving layer. It sits between the model and your applications, accepts requests over HTTP, and schedules the GPU work. Its scheduler uses continuous batching. New requests join a running batch as soon as there is room, instead of waiting for the slowest request to finish. The GPU stays busy, and short requests do not get stuck behind long ones.
The feature SGLang is best known for is RadixAttention. While generating text, a model builds a KV cache that stores intermediate results for every token it has already processed. A basic setup recomputes that work for each new request, even when thousands of requests start with the same system prompt or the same document. RadixAttention stores the cache in a tree and matches incoming prompts against it. Shared prefixes are computed once and reused. Chat assistants with long system prompts, retrieval-augmented generation, and multi-step agents gain the most from this.
Structured output is the second strength. SGLang can constrain generation to a JSON schema, a regular expression, or a grammar. Pipelines that parse model responses, such as data extraction or tool calling, receive valid output without retry loops or fragile post-processing.
The runtime covers more ground than these two features. It supports quantized models, tensor parallelism across several GPUs, speculative decoding, vision language models, and embedding models. Popular open-weight families such as Llama, Qwen, and DeepSeek run on it. The HTTP server speaks an OpenAI-compatible API, so an application that already uses an OpenAI client library only needs a new base URL to talk to your own server.
SGLang also ships with a Python frontend for writing programs that make several model calls, with branching and parallel steps. Most self-hosted deployments skip it and use only the server, which is a sensible way to start.
In practical terms, the difference shows up in three places. Throughput rises because the GPU handles many requests in parallel. Latency drops for repeated prompts because cached prefixes are skipped. Operations get easier because the server provides health checks, metrics, and a stable API, in place of a custom script that each team maintains on its own.
A serving layer is not free, though. It reserves memory for the cache and needs some configuration, such as how much VRAM to allocate and which context length to allow. The defaults work for a first run, and tuning can follow once real traffic arrives.
Installing SGLang with Docker
Docker is the most direct way to run SGLang. The official SGLang Docker image bundles the inference engine, the CUDA libraries, and the optimized kernels, so there is no need to match Python, PyTorch, and CUDA versions by hand. The host needs a working NVIDIA driver, Docker, and the NVIDIA Container Toolkit, which lets containers access the GPU. On a Contabo GPU VPS, the driver, CUDA, and the toolkit come preinstalled with the Ubuntu 24.04 image. Docker itself can be added from the 1-Click App library or installed manually.
Before pulling anything, confirm that the GPU is visible on the host and inside a container.
nvidia-smi
docker run --rm --gpus all ubuntu nvidia-smiBoth commands should print the same GPU table. The CUDA version in the top right corner of the output is worth noting, because it decides which image tag fits.
Step 1: Pull the SGLang Docker image
The images are published on Docker Hub as lmsysorg/sglang. The runtime variant is about 40 percent smaller than the full image and is the better choice for a production server.
docker pull lmsysorg/sglang:latest-runtimeThe latest tags move whenever a new build is released. For a stable setup, browse the tag list on Docker Hub and pin a specific version tag instead. Hosts running CUDA 12 should use a tag with the cu129 suffix, for example lmsysorg/sglang:latest-cu129.
Step 2: Run the container and load a model
A single docker run command starts the container and the SGLang server in one go. The example below serves Qwen2.5-7B-Instruct, a small model that needs no Hugging Face approval and fits on almost any GPU.
docker run -d \
--name sglang \
--gpus all \
--shm-size 32g \
--ipc=host \
--restart unless-stopped \
-p 127.0.0.1:30000:30000 \
-v ~/.cache/huggingface:/root/.cache/huggingface \
--env "HF_TOKEN=<your-token>" \
lmsysorg/sglang:latest-runtime \
python3 -m sglang.launch_server \
--model-path Qwen/Qwen2.5-7B-Instruct \
--host 0.0.0.0 \
--port 30000Each part of the command has a job. The gpus flag hands the GPU to the container. The shm-size and ipc flags give the server enough shared memory for its worker processes. The volume mount keeps downloaded weights on the host, so a restart does not trigger a new download. The port mapping publishes the API on the local interface only, which keeps it away from the public internet until a reverse proxy is in place. Inside the container, the host flag set to 0.0.0.0 makes the server listen on all interfaces, so the port mapping can reach it.
The HF_TOKEN line is only required for gated models, such as the Llama family. Replace the placeholder with a token from the Hugging Face account settings, or remove the line for open models. To serve a different model, change the model-path value to any Hugging Face repository that fits into the available VRAM.
Step 3: Watch the startup
The first start downloads the weights and prepares the kernels, which takes several minutes. Follow the log to see the progress.
docker logs -f sglangThe server is ready once the log reports that it is up and listening. Press Ctrl+C to leave the log view. The container keeps running in the background.
Step 4: Verify the endpoint
A health check confirms that the server responds.
curl http://localhost:30000/healthThe next command lists the loaded model through the OpenAI-compatible API.
curl http://localhost:30000/v1/modelsA test request to the chat completions endpoint confirms that the model generates text.
curl http://localhost:30000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "Qwen/Qwen2.5-7B-Instruct",
"messages": [
{"role": "user", "content": "Explain what a KV cache is in one sentence."}
],
"max_tokens": 100
}'A JSON response with a generated answer means the GPU, the model, and the API all work. The same server also provides interactive API documentation at http://localhost:30000/docs.
Step 5: Move to Docker Compose
A Compose file replaces the long docker run command with a configuration that is easy to read, version, and update. It also adds an API key, which the plain command above lacks. Create a file named docker-compose.yml with the following content.
services:
sglang:
image: lmsysorg/sglang:latest-runtime
container_name: sglang
restart: unless-stopped
ipc: host
shm_size: 32g
ports:
- "127.0.0.1:30000:30000"
volumes:
- hf-cache:/root/.cache/huggingface
environment:
- HF_TOKEN=${HF_TOKEN}
command: >
python3 -m sglang.launch_server
--model-path Qwen/Qwen2.5-7B-Instruct
--host 0.0.0.0
--port 30000
--api-key ${SGLANG_API_KEY}
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: all
capabilities: [gpu]
volumes:
hf-cache:Store the two variables in a file named .env next to the Compose file.
HF_TOKEN=<your-token>
SGLANG_API_KEY=<a-long-random-string>Stop the earlier container, then start the stack.
docker rm -f sglang
docker compose up -dWith an API key set, every request to the API needs an Authorization header with the key as a bearer token.
curl http://localhost:30000/v1/chat/completions \
-H "Authorization: Bearer <a-long-random-string>" \
-H "Content-Type: application/json" \
-d '{"model": "Qwen/Qwen2.5-7B-Instruct", "messages": [{"role": "user", "content": "Hello"}]}'Step 6: Expose the API with HTTPS
To reach the server from other machines, place a reverse proxy in front of it. Caddy is a good fit because it requests and renews TLS certificates automatically. Point a DNS record at the server, then add this block to the Caddyfile.
llm.example.com {
reverse_proxy 127.0.0.1:30000
}After reloading Caddy, applications can use https://llm.example.com/v1 as the base URL of any OpenAI client library, together with the API key from the .env file. Keep the firewall closed for port 30000 and open only ports 80 and 443.
Updating the server later is a short routine. Change the image tag in the Compose file, run the pull command for the stack, and restart it.
docker compose pull
docker compose up -dWhen SGLang Is the Right Choice
SGLang pays off when many requests share the same GPU. Typical examples include an internal assistant used by a whole team, a customer-facing chatbot, a retrieval-augmented generation service where every request carries the same instructions and similar context, and agent systems that call the model dozens of times per task. Batch jobs, such as classifying or summarizing thousands of documents overnight, also fit well, because continuous batching keeps the GPU saturated. Services that need guaranteed JSON output are another strong match.
It is a weaker choice in some cases. A single person chatting with a model a few times per day gains little from advanced scheduling, and a simpler tool such as Ollama is easier to run. Models too large for one server, which require several machines working together, call for a cluster setup that goes beyond a single VPS. Teams that cannot manage a server at all may prefer a hosted API.
Hardware planning starts with VRAM. The model weights must fit in GPU memory, and the KV cache needs room on top of them. A rough estimate for the weights is the parameter count multiplied by the bytes per parameter. A 70 billion parameter model needs about 140 GB at 16-bit precision, about 70 GB at 8-bit, and about 35 GB at 4-bit. More free VRAM means a larger cache, which means more concurrent requests and longer contexts. System RAM, fast NVMe storage for loading weights, and generous traffic limits matter as well, since model files are large and responses stream continuously.
For this kind of workload, the GPU VPS from Contabo is a fitting match. Each instance has a dedicated NVIDIA RTX PRO 6000 Blackwell Server Edition with 96 GB of VRAM, passed through to the server without sharing. It comes with 18 vCPUs, 96 GB of RAM, and 900 GB of NVMe storage. The server runs Ubuntu 24.04 with CUDA and the NVIDIA Container Toolkit already installed, so the SGLang Docker steps above work right after the first SSH login. A 70B-class model fits at 8-bit precision, and a mid-sized model leaves a very large cache for many parallel users.
Pricing is a flat monthly rate that starts at €999 excluding VAT, with discounts for longer terms. Traffic is unlimited under a fair use policy, and the product is available in the EU and US Central regions. Check the current pricing page before ordering, because capacity and regions are still expanding. Workloads that outgrow a single card can move to the multi-GPU servers offered by vshosting, a sister company of Contabo.
FAQ: Self-Hosting SGLang
Is SGLang free to self-host?
Yes, the software is free. SGLang is open source under the Apache 2.0 license, so there are no license fees for running it on your own server, including for commercial projects. The real costs sit elsewhere. A GPU server is the main expense, followed by storage, bandwidth, and the time spent on setup and maintenance. The model matters too. Many open-weight models use permissive licenses, while others, including the Llama family, come with their own terms that you should read before commercial use.
Does SGLang need a GPU?
For practical use, yes. SGLang is built for GPU inference, and NVIDIA GPUs with CUDA are the best-supported path. The project also supports AMD GPUs and several other accelerators, though maturity varies by hardware and model. Running on a CPU alone is rarely a realistic option beyond testing, because response times for modern models become too slow. When choosing a card, VRAM matters more than raw compute, since the model and its cache have to fit in memory.
How is SGLang different from a basic model server?
A basic model server loads the weights and generates text, usually handling requests one after another or in fixed batches. SGLang adds the scheduling and caching that a basic server lacks. Continuous batching keeps the GPU busy under load. RadixAttention reuses the KV cache across requests with shared prefixes. Constrained decoding enforces JSON or regex output. Multi-GPU parallelism, quantization, and an OpenAI-compatible API come built in. For a single user with light traffic, the difference is small. With many concurrent users or repeated long prompts, SGLang delivers noticeably higher throughput and lower latency on the same hardware.

