Если вы раньше входили через Google, сбросьте пароль для своей Gmail-почты через кнопку «Забыли пароль?» на экране входа. Затем войдите по email и новому паролю.
Если аккаунта ещё нет, зарегистрируйтесь с Gmail-почтой — после подтверждения почты мы предложим задать пароль.
Что нового
Загружаю обновления...
Что нового
Загружаю обновления...
Работа найдется быстрее с подпискойКандидат найдётся быстрее с подпиской
Чтобы адаптировать резюме под вакансию или составить сопроводительное письмо, загрузите резюме
описание
EPAM provides enterprise software products, open source solutions, and accelerators.
задачи
Collaborate with development, security, quality, and operations teams to implement SRE practices and ensure system reliability;
Define and support the required level of reliability, availability, and performance for services and applications;
Troubleshoot, mitigate, and support fixing infrastructure and application issues in a timely manner;
Implement monitoring systems for infrastructure and application reliability;
Define and track Service Level Objectives (SLOs);
Manage error budgets;
Reduce toil through automation;
Deploy, maintain, and automate infrastructure and application environments;
Enhance system reliability and optimize resource utilization;
Drive technology initiatives and maximize their impact across the organization.
требования
Bachelor’s degree in Computer Science, Engineering, or a related field;
Proven experience with any cloud platform: AWS, GCP, or Azure;
Experience implementing SRE practices, including SLO/SLI, error budgets, postmortems, reducing toil, capacity planning, and incident management;
Knowledge of Python or another scripting/programming language;
Strong background in monitoring tools;
Proficiency in CI/CD tools, infrastructure as code, and configuration management;
Solid knowledge of container orchestration technologies, including Kubernetes and Docker;
Nice to have: Expertise in deployment and management of LLMs, including RAG; certification in Kubernetes, AWS/GCP/Azure, or similar technologies; proven experience in DevOps; knowledge of managing and optimizing AI/ML models in production environments, including basic deployment, monitoring, and maintenance.