сегодня

site reliability engineer for customer-facing systems

ориентир по рынку
вакансия зп не указана
в среднем 325 071 ₽
Загрузи резюме, чтобы видеть мэтчи с вакансией

подготовьтесь к отклику

ai-инструменты

Чтобы адаптировать резюме под вакансию или составить сопроводительное письмо, загрузите резюме

описание

EPAM develops enterprise software products, open source solutions, and accelerators.

задачи

  • Own the observability charter for the platform by building monitoring, alerting, synthetic checks, dashboards, and runbooks;
  • Define meaningful SLIs/SLOs and reduce alert noise to improve signal quality;
  • Design and optimize release pipelines with progressive delivery, health gates, and automated rollback mechanisms;
  • Apply a performance engineering mindset through load testing, capacity analysis, and latency profiling;
  • Automate operational toil through scripting and infrastructure-as-code;
  • Accelerate SRE maturity by applying AIOps capabilities to improve detection, diagnosis, and reduce manual effort;
  • Lead incident response practices, including on-call readiness and blameless post-mortems;
  • Collaborate with DevOps, Cloud teams, product engineering teams, and Tech Leads to drive reliability improvements.

требования

  • 3+ Years of experience in Site Reliability Engineering, DevOps, or platform/production engineering supporting customer-facing systems;
  • Expertise in observability tools such as Grafana, Prometheus, and log/trace aggregation with Loki, Tempo, and OpenTelemetry, covering metrics, logs, traces, and events;
  • Knowledge of SRE fundamentals: SLIs/SLOs, error budgets, golden signals, alert tuning and noise reduction, and blameless post-incident reviews;
  • Experience operating workloads on Kubernetes, ideally EKS, and AWS, with the ability to debug issues across application, container, and infrastructure layers;
  • Proficiency in Python, Bash, or Go, with exposure to infrastructure-as-code using Terraform and CI/CD pipelines;
  • Background in incident management, including triage, escalation, communication, post-mortems, and on-call processes and rotations;
  • A proactive ownership mindset with the ability to identify problems from telemetry before they are reported and follow through with engineering teams;
  • Strong communication skills to turn noisy signals into crisp findings, runbooks, and recommendations;
  • English at B2+ (Upper-Intermediate) level or higher;
  • Nice to have: experience setting up synthetic monitoring, performance engineering skills including load/stress testing with k6, JMeter, or Locust, capacity planning and latency profiling, familiarity with AIOps, experience with Datadog or similar enterprise observability platforms, experience evangelizing best practices and setting standards across engineering teams, exposure to programmatic advertising or adtech platforms.

условия

  • Location-specific conditions and benefits are available.

Если просят выйти из iCloud, прислать код из SMS, запустить или установить что-то, перевести деньги — не соглашайтесь: это мошенничество.

Про зарплаты

Анонимные данные по зарплатам и грейдам.
Можно сверить вилку с рынком.

Посмотреть зарплаты

Если просят выйти из iCloud, прислать код из SMS, запустить или установить что-то, перевести деньги — не соглашайтесь: это мошенничество.