platform engineer for AI infrastructure
ориентир по рынку
вакансия
зп не указана
в среднем
320 875 ₽
мэтч
Загрузи резюме, чтобы видеть мэтчи с вакансией
Загрузить
Чтобы адаптировать резюме под вакансию или составить сопроводительное письмо, загрузите резюме
загрузить резюме
описание
The company provides AI infrastructure for large-scale model training and high-volume inference.
задачи
Build and operate GPU-enabled Kubernetes infrastructure; Support distributed model training and large-scale inference workloads; Optimise GPU scheduling, allocation, utilisation, and performance; Build distributed compute environments using Ray; Develop platform tooling and automation in Python; Provision infrastructure using Terraform; Improve observability across GPU, Kubernetes, and workload layers; Troubleshoot performance bottlenecks across compute, memory, networking, and storage; Build self-service capabilities for ML and AI engineering teams; Improve platform reliability, scalability, and compute efficiency.
требования
4+ Years in ML Infrastructure, Platform Engineering, MLOps, HPC, or similar roles; Kubernetes; NVIDIA GPU infrastructure; CUDA ecosystem; Python; Terraform; Ray or comparable distributed compute technologies; Monitoring and observability; Strong understanding of Linux and distributed systems; Nice to have: NVIDIA GPU Operator, NCCL and distributed GPU communication, PyTorch distributed training, Triton Inference Server / vLLM, Prometheus / Grafana / OpenTelemetry, AWS, GCP, or Azure GPU infrastructure, Slurm or HPC environments, GPU capacity planning and cost optimisation, large-scale LLM training or inference.
условия
Remote work across the European Union.
Если просят выйти из iCloud, прислать код из SMS, запустить или установить что-то, перевести деньги — не соглашайтесь: это мошенничество.
Про зарплаты
Анонимные данные по зарплатам и грейдам. Можно сверить вилку с рынком.
Посмотреть зарплаты
статьи для DevOps-инженеров