Egor Andreev
Vprod
An ML team has come to you: "The LLM is responding in the container, now give it to the users and maintain it."
Let's explore what can go wrong on the way to production: weights and GPUs, the first token, keys and accesses, updates, metrics, and crashes. And how to handle it with minimal losses.
Vprod