Если вы раньше входили через Google, сбросьте пароль для своей Gmail-почты через кнопку «Забыли пароль?» на экране входа. Затем войдите по email и новому паролю.
Если аккаунта ещё нет, зарегистрируйтесь с Gmail-почтой — после подтверждения почты мы предложим задать пароль.
Что нового
Загружаю обновления...
Что нового
Загружаю обновления...
Работа найдется быстрее с подпискойКандидат найдётся быстрее с подпиской
Чтобы адаптировать резюме под вакансию или составить сопроводительное письмо, загрузите резюме
описание
EPAM provides enterprise software products, open source solutions, and accelerators.
задачи
Collaborate with development, security, quality, and operations teams to implement SRE practices and ensure system reliability;
Define and support the required level of reliability, availability, and performance for services and applications;
Troubleshoot, mitigate, and support fixing infrastructure and application issues in a timely manner;
Implement monitoring systems for infrastructure and application reliability;
Define and track Service Level Objectives (SLOs);
Manage error budgets;
Reduce toil through automation;
Deploy, maintain, and automate infrastructure and application environments;
Optimize resource utilization and continuously improve operational practices;
Drive technology initiatives and maximize their impact across the organization;
Ensure solutions meet customer expectations and required standards.
требования
Bachelor’s degree in Computer Science, Engineering, or a related field;
Proven experience with any cloud platform, including AWS, GCP, or Azure;
Experience implementing SRE practices such as SLO/SLI, error budgets, postmortems, reducing toil, capacity planning, and incident management;
Knowledge of Python or another scripting or programming language;
Strong background in monitoring tools;
Proficiency in CI/CD tools, infrastructure as code, and configuration management;
Solid knowledge of container orchestration technologies, including Kubernetes and Docker;
Nice to have: expertise in deployment and management of LLMs, including RAG; certification in Kubernetes, AWS/GCP/Azure, or similar technologies; proven experience in DevOps; knowledge of managing and optimizing AI/ML models in production environments, including basic deployment, monitoring, and maintenance.