27 сен

ML infrastructure engineer for AI cloud

ориентир по рынку
вакансия зп не указана
в среднем 328 556 ₽
Загрузи резюме, чтобы видеть мэтчи с вакансией

подготовься к отклику

ai-инструменты

Чтобы адаптировать резюме под вакансию или составить сопроводительное письмо, загрузи резюме

Рекламный баннер: ОБЩЕСТВО С ОГРАНИЧЕННОЙ ОТВЕТСТВЕННОСТЬЮ "ЦЕНТР НАЦИОНАЛЬНЫХ ИНТЕЛЛЕКТУАЛЬНЫХ СИСТЕМ"
О рекламодателе
ОБЩЕСТВО С ОГРАНИЧЕННОЙ ОТВЕТСТВЕННОСТЬЮ "ЦЕНТР НАЦИОНАЛЬНЫХ ИНТЕЛЛЕКТУАЛЬНЫХ СИСТЕМ"
ИНН: 9704271170

описание

Nebius builds a full-stack AI cloud platform that supports developers and enterprises from data and model training through production deployment. Its infrastructure spans compute, storage, networking, and applied AI.

задачи

  • Profile and analyze GPU performance at the system and kernel level with hardware and development teams
  • Evaluate and compare GPU performance across platforms, architectures, and software stacks such as CUDA and ROCm
  • Debug and optimize ML workloads for GPU hardware, identifying and resolving performance bottlenecks
  • Perform acceptance testing for new GPU clusters to ensure hardware and software meet performance, stability, and compatibility requirements for AI workloads
  • Run experiments across GPU system configurations to assess how interconnect strategies and system-level optimizations affect performance and scalability
  • Develop tools and dashboards to visualize performance metrics, bottlenecks, and trends
  • Contribute to internal tooling, frameworks, and best practices

требования

  • Strong understanding of the theoretical foundations of machine learning
  • Deep understanding of performance aspects of large neural network training and inference, including parallelism, offloading, custom kernels, hardware features, attention optimizations, and dynamic batching
  • Deep experience with modern deep learning frameworks, including PyTorch, JAX, Megatron-LM, and Tensort-LLM
  • Good understanding of the GPU stack, including CUDA, NCCL, drivers, and relevant libraries
  • Familiarity with containerized environments such as Docker and Kubernetes
  • Strong communication skills and ability to work independently
  • Будет плюсом: familiarity with modern LLM inference frameworks such as vLLM, SGLang, and TensorRT, experience with Python and performance profiling tools such as Nsight, nvprof, and perf, familiarity with cloud ML platforms such as AWS, GCP, and Azure ML, contributions to open-source ML benchmarking tools

условия

  • Competitive compensation
  • Career growth and learning opportunities
  • Flexibility and ownership
  • Collaborative and innovative culture
  • Opportunity to work on impactful AI projects
  • International environment and talented teams
  • Applicants must be authorized to work in the country in which they apply and provide proof of employment eligibility as a condition of hire

ИИ-облако для обучения и запуска моделей.

Если просят выйти из iCloud, прислать код из SMS, запустить или установить что-то, перевести деньги — не соглашайся: это мошенничество.

Про зарплаты

Анонимные данные по зарплатам и грейдам.
Можно сверить вилку с рынком.

Посмотреть зарплаты

Если просят выйти из iCloud, прислать код из SMS, запустить или установить что-то, перевести деньги — не соглашайся: это мошенничество.