Codea Bien Logo
vLLM from zero: your first inference server

vLLM from zero: your first inference server

From zero to your own vLLM server on a rented GPU: 8 hands-on chapters, for under $5, with measured numbers at every step.

8 articles

A hands-on guide to go from zero inference knowledge to your own vLLM server on a rented GPU: for under $5, with measured numbers at every step.

What you'll build

  • A local model running on your machine with Ollama

  • A vLLM server with a 30B model in FP8 on a 48 GB GPU

  • A gateway with per-person keys, budgets, and logs

  • Your own TTFT, tokens/s, and concurrency benchmarks

8 chapters, each with commands that ran on a real pod and a time/cost checkpoint. Start at chapter 1 and don't skip steps: each one builds on the previous.

Guide contents

  1. 01

    What Serving an AI Model Actually Means (and Why vLLM Exists)

  2. 02

    Your First Local Model in 10 Minutes (Ollama)

  3. 03

    Rent Your First GPU Without Fear (~$1 per hour)

  4. 04

    Your First vLLM Server (hello world)

  5. 05

    The 5 vLLM Flags That Actually Matter

  6. 06

    Measure Your vLLM: TTFT, tokens/s and Concurrency Without Fooling Yourself

  7. 07

    LiteLLM: Keys, Budgets and Your vLLM's Bill

  8. 08

    When It Explodes: The 7 Errors of Serving vLLM on a Rented GPU