Codea Bien Logo
</>

My First Self-Hosted Inference Server

The complete PoC: what I measured, what I broke and what it cost to serve an open-source model on a rented GPU.

8 articles

Serving your own model stopped being a lab topic: today you can rent a GPU for under a dollar an hour and an open 30B model fits in 48 GB of VRAM. The question is no longer is it possible, but is it worth it for you.

This series is the complete PoC we ran to answer that with numbers: two engineers, one rented L40S, a MoE code model and one rule — nothing gets decided without measuring. Eight chapters with the stack (opencode -> LiteLLM -> vLLM), the hardware and its real costs, the bugs that cost us hours, the latency and throughput measurements, the memory math and the final report that —spoiler— recommended not migrating yet, math in hand.

You'll find: setups you can copy, the 7 bugs of a rented environment with symptom, cause and fix, and real numbers of our own (TTFT, tokens/s, VRAM and cost) instead of someone else's benchmarks. And also what we did not measure: model quality, private-network latency and behavior with 30 concurrent users. That honesty is part of the value: a PoC that says "not yet" saves you an expensive decision.

Start with chapter 1 (the plan) and follow the order: the series is meant as a journey, from "can we?" to "now what?".

Guide contents

  1. 01

    Plan: serving my own model to cut the bill

  2. 02

    The 7 bugs of a rented GPU environment

  3. 03

    The error in my own benchmark

  4. 04

    Prefix caching: same load, 3x faster

  5. 05

    The bottleneck was not the GPU: it was the network

  6. 06

    48 KB per token: why 48 GB cannot hold 30 sessions

  7. 07

    Inside MoE: 30B total, 3B active

  8. 08

    The report that said "do not migrate yet"