</>
Inference Engineering: Zero To Production
Correr modelos de verdad: local, cuantización, VRAM, serving, medición y costos.
8 articles
Guide contents
- 01
Inference Engineering: what it is and what you optimize (TTFT, tokens/s, VRAM, cost)
- 02
Local inference in 2026: from LM Studio to vLLM (with VRAM math)
- 03
Quantization with numbers: quality vs VRAM vs speed
- 04
What hardware do you need? GPU, VRAM and cloud
- 05
Serving in production: batching, concurrency and autoscaling
- 06
Measuring inference without lying to yourself: tokens/s, TTFT and misleading benchmarks
- 07
The real cost of your AI feature: API vs self-host vs edge
- 08
Serving multiple models in production: routing, versioning and canary