Если вы раньше входили через Google, сбросьте пароль для своей Gmail-почты через кнопку «Забыли пароль?» на экране входа. Затем войдите по email и новому паролю.
Если аккаунта ещё нет, зарегистрируйтесь с Gmail-почтой, после подтверждения почты мы предложим задать пароль.
Что нового
Загружаю обновления...
Что нового
Загружаю обновления...
Работа найдется быстрее с подпискойКандидат найдётся быстрее с подпиской
infrastructure engineer in global cloud infrastructure
ориентир по рынку
вакансия
зп не указана
в среднем
320 875 ₽
мэтч
Загрузи резюме, чтобы видеть мэтчи с вакансией
подготовьтесь к отклику
ai-инструменты
Чтобы адаптировать резюме под вакансию или составить сопроводительное письмо, загрузите резюме
описание
EPAM develops enterprise software products, open source solutions, and technology accelerators through its SolutionsHub and related brands.
задачи
Design, build, and maintain a global, distributed, and resilient cloud infrastructure;
Collaborate with infrastructure and product engineering teams to plan and deliver complex platform initiatives;
Participate in architecture reviews, incident response, and performance analysis to ensure system reliability;
Manage and provision AWS infrastructure using Terraform and Kubernetes;
Write and maintain Kubernetes manifests and deployment configurations for critical workloads, including pod security contexts, anti-affinity rules, network policies, autoscaling, and health probes;
Drive Production Readiness Reviews for all new services, covering security, HA, performance, and observability gates;
Design and operate multi-layer workload isolation using Linux kernel primitives, including namespaces, cgroups, seccomp profiles, and capabilities;
Evaluate and operate gVisor and Firecracker for workloads requiring hard tenant boundaries and near-native performance;
Design and maintain Grafana dashboards;
Manage Prometheus and VictoriaMetrics pipelines; define and tune P1/P2/P3 alert thresholds with runbooks;
Contribute to and extend the internal k6-based load testing framework;
Design load scenarios using constant-arrival-rate profiles and instrument custom metrics tagged by testid for Grafana correlation;
Stream test metrics to Prometheus via remote write;
Generate and publish HTML reports to file storage after each run.
требования
5+ Years of experience in infrastructure, SRE, or platform engineering roles;
Expertise in distributed systems and cloud-native architectures;
Understanding of Linux internals, including namespaces, cgroups, seccomp, capabilities, and system-level performance tuning;
Experience operating infrastructure on AWS at scale;
Proficiency in Terraform and Kubernetes, including security hardening of manifests;
Experience designing and running load tests with k6, Gatling, Locust, or similar tools;
Understanding of network security and cloud security best practices;
Excellent analytical, troubleshooting, and communication skills;
English proficiency at a B2+ level;
Nice to have: Skills in Golang/Python for building internal tooling, familiarity with Kafka, Redis, ClickHouse, or PostgreSQL, familiarity with Grafana, Prometheus, or VictoriaMetrics, knowledge of CMK/DEK/KMS encryption key hierarchies and HashiCorp Vault, hands-on production experience with gVisor or Firecracker, prior work extending or maintaining an internal testing framework, contributions to or maintenance of open-source infrastructure projects.
условия
Location-specific conditions and benefits are available.