сегодня

devops engineer in HPC and MLOps

ориентир по рынку
вакансия зп не указана
в среднем 323 895 ₽
Загрузи резюме, чтобы видеть мэтчи с вакансией

подготовься к отклику

ai-инструменты

Чтобы адаптировать резюме под вакансию или составить сопроводительное письмо, загрузи резюме

Рекламный баннер: ОБЩЕСТВО С ОГРАНИЧЕННОЙ ОТВЕТСТВЕННОСТЬЮ "ЦЕНТР НАЦИОНАЛЬНЫХ ИНТЕЛЛЕКТУАЛЬНЫХ СИСТЕМ"
О рекламодателе
ОБЩЕСТВО С ОГРАНИЧЕННОЙ ОТВЕТСТВЕННОСТЬЮ "ЦЕНТР НАЦИОНАЛЬНЫХ ИНТЕЛЛЕКТУАЛЬНЫХ СИСТЕМ"
ИНН: 9704271170

описание

The Science, Innovation & Labs team is responsible for scaling, reliability, and automation of a high-performance computing (HPC) and machine learning operations (MLOps) platform.

задачи

  • Guide scientists and data teams in using the platform UI effectively and running self-service workloads without infrastructure friction
  • Advise users and manage infrastructure capacity, optimizing costs, quotas, and resource availability for heavy workloads
  • Maintain automated pipelines for infrastructure provisioning and platform service deployments
  • Resolve technical queries about job scheduling failures, cluster bottlenecks, and resource quotas
  • Collaborate with developer experience teams to improve documentation
  • Collaborate with engineering teams to monitor GPU utilization using tools such as CloudWatch or Prometheus
  • Manage AWS GPU instance families and allocate block compute for large-scale ML training and inference pipelines
  • Ensure compute availability through capacity planning and reservation management
  • Deploy containerized environments tuned for HPC and GPU pass-through
  • Deploy and scale HPC workloads on cloud infrastructure using parallel storage and networking solutions

требования

  • 5+ Years of experience in HPC or DevOps engineering roles
  • Knowledge of MPI, OpenMP, and multi-node GPU communication protocols such as NCCL and GPUDirect
  • Proven experience managing AWS GPU instance families, including P-series, G-series, and Tranium/Inferentia
  • Hands-on mastery of AWS Capacity Blocks for ML, On-Demand Capacity Reservations (ODCRs), and Service Quota management
  • Experience deploying containerized environments using Apptainer/Singularity, Docker, or Enroot
  • Understanding of I/O performance bottlenecks when interfacing with distributed file systems such as Lustre, GPFS, BeeGFS, or AWS FSx for Lustre
  • Hands-on skill in profiling applications using NVIDIA Nsight or similar tools to locate memory and compute bottlenecks
  • Experience deploying or scaling HPC workloads on cloud infrastructure utilizing EFA, ParallelCluster, and parallel storage (FSx for Lustre)
  • English proficiency at B2+ level

условия

  • Условий в вакансии нет

Глобальная компания в сфере digital engineering, product development и технологического консалтинга.

Если просят выйти из iCloud, прислать код из SMS, запустить или установить что-то, перевести деньги — не соглашайся: это мошенничество.

Про зарплаты

Анонимные данные по зарплатам и грейдам.
Можно сверить вилку с рынком.

Посмотреть зарплаты

Если просят выйти из iCloud, прислать код из SMS, запустить или установить что-то, перевести деньги — не соглашайся: это мошенничество.