Talk

Weights, GPU, and the First Token: How to Launch an LLM

In Russian

An ML team has come to you: "The LLM is responding in the container, now give it to the users and maintain it."

Let's explore what can go wrong on the way to production: weights and GPUs, the first token, keys and accesses, updates, metrics, and crashes. And how to handle it with minimal losses.

Speakers

Schedule