Если вы раньше входили через Google, сбросьте пароль для своей Gmail-почты через кнопку «Забыли пароль?» на экране входа. Затем войдите по email и новому паролю.
Если аккаунта ещё нет, зарегистрируйтесь с Gmail-почтой, после подтверждения почты мы предложим задать пароль.
Что нового
Загружаю обновления...
Что нового
Загружаю обновления...
Работа найдется быстрее с подпискойКандидат найдётся быстрее с подпиской
Чтобы адаптировать резюме под вакансию или составить сопроводительное письмо, загрузите резюме
описание
EPAM is a global provider of digital engineering, cloud, and AI-enabled transformation services, focusing on complex software product development and digital platform engineering.
задачи
Lead and scale a global SRE organization, focusing on engineering excellence and team empowerment;
Collaborate with product, platform, operations, and security teams to embed reliability within SDLC practices;
Define and monitor KPIs for system reliability, performance, and operational efficiency;
Advance automation, Infrastructure as Code approaches, and promote self-healing systems using AI/ML techniques;
Develop robust incident management frameworks and lead major incident response activities for critical systems;
Implement blameless postmortems and deliver systemic improvements across production environments;
Establish observability strategies with standardized tooling for metrics, logs, and tracing to support distributed systems;
Adopt and enforce SRE practices, including SLIs, SLOs, SLAs, and error budgets across services;
Drive resilience strategies with highly available architectures and disaster recovery readiness;
Champion an automation-first culture, leveraging CI/CD pipelines and operational tooling to reduce manual processes.
требования
Strong background in Site Reliability Engineering, DevOps, or platform operations in complex, distributed environments;
Expertise in observability platforms, troubleshooting distributed systems, and telemetry-driven insights;
Hands-on experience with automation, Infrastructure as Code (Terraform or CloudFormation), and CI/CD practices;
Deep understanding of incident management processes, ITSM standards, and ITIL principles;
Knowledge of resilience design patterns, high availability, and fault-tolerant architectures;
Familiarity with AI/ML-driven approaches for operational efficiency and system reliability;
Ability to lead transformation, influence across teams, and foster continuous improvement in culture;
Nice to have: Experience in financial services or other highly regulated, mission-critical environments, Certifications in cloud technologies such as AWS, Exposure to AIOps platforms or advanced observability tooling.