Opening of the Conference
We'll talk about the schedule, sessions and share information.
New talks are published weekly. Follow updates or secure your ticket early.
We'll talk about the schedule, sessions and share information.
The talk examines the path from AI-powered operation to self-managed infrastructure: levels of autonomy, agent-based control loops, dynamic sandboxes, operational security, verification, and the changing role of the human operator.
MWS Cloud Platform
The story of DORA metrics implementation: from decomposition and value formulation to implementation results and error analysis.
Kontur
Hardening is not just something you deploy once: you need to continuously verify that it is applied across all hosts and has not drifted over time. In this talk, I will explain how we fully automated this process, from building the source of truth and dynamic inventory to running Ansible checks and visualizing the results with Prometheus and Grafana.
MWS Cloud Platform
I'll tell you how we migrated Kafka from VM to Kubernetes.
Mindbox
I will explain how the Middleware Service Hub transforms disparate data into a connected technology context, evolves from a component catalog to an organization's Digital Twin, and becomes the foundation for Architecture as Code, API governance, and AI Harness.
Raiffeisen Bank
How our fault-tolerant K8s clusters are structured in the infrastructure, and how we have automated their management through CAPI on a scale of hundreds of servers.
MWS Cloud Platform
I'll show you how to build a cascade deployment at FluxCD and not to spread the erroneous change to all environments at once.
Mindbox
A talk on building a large "hard" infrastructure and how we at K2 Cloud have gone from complete chaos to a standardized yet flexible system that allows us to quickly adapt to the dynamic pace of cloud platform development.
K2 Cloud
How to build a scalable system of dynamic test environments and avoid the problems associated with shared test benches.
Blackhub Games
The talk provides practical experience on scaling, observability, and awareness of policies, so that each team understands what is happening in their cluster without waiting for an incident.
MWS Cloud Platform
The hottest hiring trends questions from big-tech interviewers. Come to feel the engineering culture and refresh your knowledge!
judies.tech
If you want to feel like a DevVibeOps, then come to me! :)
Collabo
The talk is divided into two parts: network isolation of infected infrastructure and how to do it for engineers. And where to run if you need to deploy new infrastructure.
Kaspersky
Let's compare the capabilities of ArgoCD and Flux, kubectl, Helm, and Nelm/werf in organizing deployments, monitoring readiness, and managing resource lifecycles.
Flant
An ML team has come to you: "The LLM is responding in the container, now give it to the users and maintain it."
Let's explore what can go wrong on the way to production: weights and GPUs, the first token, keys and accesses, updates, metrics, and crashes. And how to handle it with minimal losses.
Vprod
We will tell you about the difficulties we faced and show you how we turned chaos into a manageable system: what went wrong at the start, which tools actually worked, and how we measured the effect.
2GIS
SRE practices and what is CRE: how to ensure the availability and fault tolerance of services (and not just infrastructure) in the cloud and on-premises.
Yandex Cloud
How to find the causes of complex problems in GPU clusters when they occur at the junction of different layers of the system and server hardware, and why the real causes of such problems are often not where you would expect to find them.
T-Bank
How the Cluster API update almost destroyed our clusters, how we mitigated it, what we changed, and how we're working now.
MWS Cloud Platform
A team workshop where you solve a real-time combat incident using board game rules.
MTS Bank
I'll describe how we created the DevOps Cookbook.
Raiffeisen Bank
Neural networks and LLMs are often perceived as impenetrable "magic," but in reality, their perimeter is less secure than that of classic IT systems.
G-HACK
I will explain how NGFW import substitution in a bank turned into a redesign of the production Kubernetes network model: namespace-level access control, cloud migration, and moving away from the familiar syslog-based authorization flow.
Tochka
The talk will cover onboarding, learning mechanisms, knowledge transfer, technologies used, and working conditions. You will learn about the perks and benefits you pay for.
RTK IT School
Let's explore popular open-source tools that can enhance Kubernetes security at runtime.
Luntry
Every other incoming resume lies (17 years of Kubernetes for a thing that is 12 years old), and the model that catches it either bankrupts you on bills, lies itself, or makes a hiring call you cannot later explain to lawyers. I show how to fix this with a three-tier model cascade with the expensive one last, three guard layers, and real before-and-after numbers — on a live collection of caught fabrications.
I will show you how to make running Spark tasks convenient for DE/ML engineers and maintain cluster management.
MAGNIT TECH
I will tell you how we built MagnitGPT, what problems we solved during development, and how it helps hundreds and thousands of users every day.
MAGNIT TECH
Let's go through two exercises that helped me develop observability.
Yandex
Real experience in evolving the approach to managing centralized configuration of Gitlab, Hashicorp Vault, and other projects in 2GIS: from simple bash scripts to a combination of Terraform and Terragrunt.
2GIS
Summing up the results of the conference, remembering the highlights and talking about plans.