Если вы раньше входили через Google, сбросьте пароль для своей Gmail-почты через кнопку «Забыли пароль?» на экране входа. Затем войдите по email и новому паролю.
Если аккаунта ещё нет, зарегистрируйтесь с Gmail-почтой — после подтверждения почты мы предложим задать пароль.
Что нового
Загружаю обновления...
Что нового
Загружаю обновления...
Работа найдется быстрее с подпискойКандидат найдётся быстрее с подпиской
Чтобы адаптировать резюме под вакансию или составить сопроводительное письмо, загрузите резюме
описание
EPAM provides enterprise software products, open source solutions, and technology accelerators.
задачи
Collaborate with development, security, quality, and operations teams to implement SRE practices and ensure system reliability;
Define and support the required level of reliability, availability, and performance for services and applications;
Troubleshoot, mitigate, and support fixing infrastructure and application issues in a timely manner;
Implement monitoring systems for infrastructure and application reliability;
Define and track Service Level Objectives (SLOs);
Manage error budgets;
Reduce toil through automation;
Deploy, maintain, and automate infrastructure and application environments;
Enhance system reliability and optimize resource utilization;
Drive technology initiatives and maximize their impact across the organization;
Ensure solutions meet customer expectations and required standards.
требования
Bachelor’s degree in Computer Science, Engineering, or a related field;
Proven experience with any cloud platform, including AWS, GCP, or Azure;
Experience implementing SRE practices such as SLO/SLI, error budgets, postmortems, reducing toil, capacity planning, and incident management;
Knowledge of Python or another scripting/programming language;
Strong background in monitoring tools;
Proficiency in CI/CD tools, infrastructure as code, and configuration management;
Solid knowledge of container orchestration technologies, including Kubernetes and Docker;
Nice to have: Expertise in deployment and management of LLMs, including RAG; certification in Kubernetes, AWS/GCP/Azure, or similar technologies; proven experience in DevOps; knowledge of managing and optimizing AI/ML models in production environments, including basic deployment, monitoring, and maintenance.