Если вы раньше входили через Google, сбросьте пароль для своей Gmail-почты через кнопку «Забыли пароль?» на экране входа. Затем войдите по email и новому паролю.
Если аккаунта ещё нет, зарегистрируйтесь с Gmail-почтой — после подтверждения почты мы предложим задать пароль.
Что нового
Загружаю обновления...
Что нового
Загружаю обновления...
Работа найдется быстрее с подпискойКандидат найдётся быстрее с подпиской
Чтобы адаптировать резюме под вакансию или составить сопроводительное письмо, загрузите резюме
описание
EPAM’s AI Workbench Platform team develops infrastructure for operationalizing domain foundation models trained on log, seismic, drilling, and production data for oil and gas subsurface and production domains. The platform enables internal teams and external customers to use, consume, and fine-tune these models at scale.
задачи
Design, build, and maintain scalable infrastructure for the AI Workbench Platform and its domain foundation models;
Implement and manage Kubernetes clusters with multi-GPU scheduling for large-scale model training and inference workloads;
Develop and maintain Infrastructure as Code using Terraform to provision cloud resources reliably and reproducibly;
Package, deploy, and manage applications with Helm and Kustomize across multiple environments;
Collaborate with data scientists, ML engineers, and business units to enable model consumption and fine-tuning workflows at scale;
Ensure platform reliability, scalability, and security for internal teams and external customers;
Optimize resource utilization and cost efficiency across GPU-intensive workloads;
Establish CI/CD pipelines and automation for platform delivery and model deployment;
Monitor system performance and troubleshoot production issues to maintain high availability;
Contribute to platform architecture decisions and MLOps best practices at enterprise scale.
требования
3+ Years of experience in DevOps, Site Reliability Engineering, or Infrastructure Engineering roles;
Expertise in Kubernetes with proven experience in multi-GPU scheduling for AI/ML workloads;
Proficiency in Terraform for Infrastructure as Code and cloud resource management;
Skills in Helm and Kustomize for Kubernetes application packaging and configuration management;
Experience building and operating production-grade platforms supporting large-scale, distributed workloads;
Understanding of MLOps principles and infrastructure requirements for training and serving large foundation models;
Ability to collaborate cross-functionally with data scientists, ML engineers, and business stakeholders;
Excellent command of written and spoken English (B2+ level);
Nice to have: Prior experience with LightOps infrastructure, familiarity with on-premises infrastructure environments, knowledge of High-Performance Computing (HPC) systems and workloads.
условия
Location-specific conditions and benefits are available depending on the selected option.