Если вы раньше входили через Google, сбросьте пароль для своей Gmail-почты через кнопку «Забыли пароль?» на экране входа. Затем войдите по email и новому паролю.
Если аккаунта ещё нет, зарегистрируйтесь с Gmail-почтой — после подтверждения почты мы предложим задать пароль.
Что нового
Загружаю обновления...
Что нового
Загружаю обновления...
Работа найдется быстрее с подпискойКандидат найдётся быстрее с подпиской
Чтобы адаптировать резюме под вакансию или составить сопроводительное письмо, загрузите резюме
описание
EPAM develops enterprise software products, open source solutions, and technology accelerators.
задачи
Implement and maintain CI/CD pipelines for AI and machine learning projects, ensuring robust deployment strategies and continuous integration;
Monitor and ensure the reliability, availability, and performance of AI applications, particularly those involving LLMs and RAG;
Collaborate with AI research teams to operationalize machine learning models and systems efficiently;
Develop and enforce best practices for version control, configuration management, and testing of AI-driven software solutions;
Use MLOps tools such as Kubeflow, MLflow, or TensorFlow Extended (TFX) to streamline the machine learning lifecycle from experimentation to production;
Implement monitoring solutions that track system metrics and model performance to facilitate proactive issue resolution;
Participate in on-call rotations to support the operational health of critical systems, applying SRE principles to meet service-level objectives (SLOs) and reduce downtime.
требования
Bachelor’s degree in Computer Science, Engineering, or a related field;
Proven experience as a DevOps Engineer or SRE, with a strong background in software development and automation;
Strong familiarity with Generative AI Operations;
Expertise in deploying and managing LLMs, including technologies such as RAG;
Proficiency with CI/CD tools, including Jenkins, GitLab CI, and CircleCI;
Proficiency with infrastructure as code using Terraform and Ansible;
Solid knowledge of container orchestration technologies, including Kubernetes and Docker;
Familiarity with MLOps tools and practices supporting machine learning lifecycle management;
Nice to have: Experience with AWS, GCP, or Azure for AI/ML deployments, Prometheus, Grafana, and ELK stack, Python in data science and machine learning contexts, certification in Kubernetes, AWS/GCP/Azure, or similar technologies.