Codea Bien Logo
</>

Inference Engineering: Zero To Production

Correr modelos de verdad: local, cuantización, VRAM, serving, medición y costos.

8 articles

Guide contents

  1. 01

    Inference Engineering: what it is and what you optimize (TTFT, tokens/s, VRAM, cost)

  2. 02

    Local inference in 2026: from LM Studio to vLLM (with VRAM math)

  3. 03

    Quantization with numbers: quality vs VRAM vs speed

  4. 04

    What hardware do you need? GPU, VRAM and cloud

  5. 05

    Serving in production: batching, concurrency and autoscaling

  6. 06

    Measuring inference without lying to yourself: tokens/s, TTFT and misleading benchmarks

  7. 07

    The real cost of your AI feature: API vs self-host vs edge

  8. 08

    Serving multiple models in production: routing, versioning and canary