vLLM from zero: your first inference server
From zero to your own vLLM server on a rented GPU: 8 hands-on chapters, for under $5, with measured numbers at every step.
8 articles
A hands-on guide to go from zero inference knowledge to your own vLLM server on a rented GPU: for under $5, with measured numbers at every step.
What you'll build
A local model running on your machine with Ollama
A vLLM server with a 30B model in FP8 on a 48 GB GPU
A gateway with per-person keys, budgets, and logs
Your own TTFT, tokens/s, and concurrency benchmarks
8 chapters, each with commands that ran on a real pod and a time/cost checkpoint. Start at chapter 1 and don't skip steps: each one builds on the previous.
Guide contents
- 01
What Serving an AI Model Actually Means (and Why vLLM Exists)
- 02
Your First Local Model in 10 Minutes (Ollama)
- 03
Rent Your First GPU Without Fear (~$1 per hour)
- 04
Your First vLLM Server (hello world)
- 05
The 5 vLLM Flags That Actually Matter
- 06
Measure Your vLLM: TTFT, tokens/s and Concurrency Without Fooling Yourself
- 07
LiteLLM: Keys, Budgets and Your vLLM's Bill
- 08
When It Explodes: The 7 Errors of Serving vLLM on a Rented GPU