Чтобы адаптировать резюме под вакансию или составить сопроводительное письмо, загрузи резюме
О рекламодателе
ОБЩЕСТВО С ОГРАНИЧЕННОЙ ОТВЕТСТВЕННОСТЬЮ "ЦЕНТР НАЦИОНАЛЬНЫХ ИНТЕЛЛЕКТУАЛЬНЫХ СИСТЕМ" ИНН: 9704271170
описание
Nebius builds a full-stack AI cloud platform that supports developers and enterprises from data and model training through production deployment. Its infrastructure spans compute, storage, networking, and applied AI.
задачи
Profile and analyze GPU performance at the system and kernel level with hardware and development teams
Evaluate and compare GPU performance across platforms, architectures, and software stacks such as CUDA and ROCm
Debug and optimize ML workloads for GPU hardware, identifying and resolving performance bottlenecks
Perform acceptance testing for new GPU clusters to ensure hardware and software meet performance, stability, and compatibility requirements for AI workloads
Run experiments across GPU system configurations to assess how interconnect strategies and system-level optimizations affect performance and scalability
Develop tools and dashboards to visualize performance metrics, bottlenecks, and trends
Contribute to internal tooling, frameworks, and best practices
требования
Strong understanding of the theoretical foundations of machine learning
Deep understanding of performance aspects of large neural network training and inference, including parallelism, offloading, custom kernels, hardware features, attention optimizations, and dynamic batching
Deep experience with modern deep learning frameworks, including PyTorch, JAX, Megatron-LM, and Tensort-LLM
Good understanding of the GPU stack, including CUDA, NCCL, drivers, and relevant libraries
Familiarity with containerized environments such as Docker and Kubernetes
Strong communication skills and ability to work independently
Будет плюсом: familiarity with modern LLM inference frameworks such as vLLM, SGLang, and TensorRT, experience with Python and performance profiling tools such as Nsight, nvprof, and perf, familiarity with cloud ML platforms such as AWS, GCP, and Azure ML, contributions to open-source ML benchmarking tools
условия
Competitive compensation
Career growth and learning opportunities
Flexibility and ownership
Collaborative and innovative culture
Opportunity to work on impactful AI projects
International environment and talented teams
Applicants must be authorized to work in the country in which they apply and provide proof of employment eligibility as a condition of hire