Если вы раньше входили через Google, сбросьте пароль для своей Gmail-почты через кнопку «Забыли пароль?» на экране входа. Затем войдите по email и новому паролю.
Если аккаунта ещё нет, зарегистрируйтесь с Gmail-почтой, после подтверждения почты мы предложим задать пароль.
Что нового
Загружаю обновления...
Что нового
Загружаю обновления...
Работа найдется быстрее с подпискойКандидат найдётся быстрее с подпиской
Чтобы адаптировать резюме под вакансию или составить сопроводительное письмо, загрузите резюме
описание
No description
задачи
Operate Kubernetes platforms (EKS, CKS, GKE) at significant scale across providers;
Own cluster lifecycle, node pool management, networking policy, and stability maintenance during rapid growth;
Provision HPC infrastructure through CI/CD systems across AWS, CoreWeave, GCP, and OCI, with additional providers to be added in the near future;
Manage job scheduling to allocate GPU compute across training and inference workloads;
Define and maintain SLIs/SLOs;
Build monitoring and alerting systems;
Participate in severity escalation response and author post-incident reviews;
Coordinate daily with Networking, Storage, Security, and AI/ML platform teams.
требования
5+ Years of experience in infrastructure engineering, cloud platforms, or HPC;
Expertise in Kubernetes, including hands-on experience operating clusters at meaningful scale, node pool sizing, scheduler debugging, CNI troubleshooting, and rolling upgrades across large fleets;
Proficiency in Terraform for writing and reviewing infrastructure-as-code daily;
Working knowledge of AWS (EC2, S3, EFS, FSx for Lustre);
Skills in Python for tooling and automation;
English proficiency at B2 level or higher;
Nice to have: Familiarity with Google Kubernetes Engine, familiarity with Amazon Elastic Kubernetes Service, knowledge of Google Cloud Platform.