сегодня

site reliability engineer for Linux security services

ориентир по рынку
вакансия зп не указана
в среднем 437 639 ₽
Загрузи резюме, чтобы видеть мэтчи с вакансией

подготовьтесь к отклику

ai-инструменты

Чтобы адаптировать резюме под вакансию или составить сопроводительное письмо, загрузите резюме

описание

Imunify360 provides a multi-layer Linux server security suite that includes WAF, IDS/IPS, malware scanning and cleanup, proactive defence, patch management, and reputation services. The platform runs as an agent on hundreds of thousands of customer servers and is supported by cloud-based scanning, correlation, and signature-delivery services on bare metal.

задачи

  • Define SLIs, SLOs, error budgets, and owning squads for approximately 70 product components;
  • Build a product-specific taxonomy covering service, fleet, control-efficacy, delivery, and pipeline indicators;
  • Enforce measurement designs that remain observable outside the gate of the measured component;
  • Design and build a push-based, sampled, privacy-constrained telemetry pipeline with a defined cardinality budget;
  • Extend agent-side and service-side instrumentation in Python, Go, and Rust;
  • Consolidate dashboards, ad-hoc queries, and reporting paths into a defensible set of instruments;
  • Build symptom-based, SLO-anchored alerting with multi-window burn-rate semantics;
  • Define page, ticket, and dashboard alert tiers and the rules for paging humans;
  • Ship alerts with owners, runbooks, and documented failure modes;
  • Establish ongoing alert hygiene, quarterly reviews, deletion tracking, and actionable-rate measurement;
  • Build a machine-readable component ownership map and connect it to alert routing;
  • Design severity matrices, acknowledgement SLAs, follow-the-sun rotations, and handoff protocols;
  • Practice incident command and blameless postmortems within 24 hours;
  • Build and operate the escalation platform and coach squads on pager ownership;
  • Deliver the first-year outcomes for component inventory, SLI adoption, collection pipeline production, alert taxonomy, escalation routing, and silent control-degradation detection;
  • Document the function so future hires can contribute without archaeological work.

требования

  • Have substantial production-engineering or SRE experience, including defining an SLO framework rather than inheriting one;
  • Be able to explain personally authored SLIs and how they were negotiated with resistant teams;
  • Have strong Python skills;
  • Be comfortable reading and modifying Go or Rust;
  • Have deep practical experience with time-series and event telemetry at scale;
  • Know Prometheus/OpenMetrics, Grafana, an Alertmanager-class routing layer, and a columnar store such as ClickHouse;
  • Have experience debugging distributed systems on bare metal and long-lived hosts;
  • Know configuration management and CI at production scale, including Ansible, GitLab CI, Jenkins, or close equivalents;
  • Be able to design measurement for machines that cannot be owned or scraped directly;
  • Have judgement regarding push telemetry, sampling, clock skew, partial reporting, and customer-server privacy constraints;
  • Have written communication skills suitable for asynchronous work;
  • Nice to have: Security product background, monitoring under audit, OpenTelemetry, eBPF, Sentry, cost- and cardinality-aware telemetry design, agentic development tooling, Kubernetes.

условия

  • Remote-first and async work across nine time zones;
  • Weekly PO and architecture syncs, monthly demo and OKR review, and a quarterly architecture summit;
  • Decisions are recorded as ADRs, and every output has an owner and due date;
  • Blameless postmortems are published, with corrections made on the record;
  • Flexible working hours;
  • Work from any location worldwide;
  • 24 Paid vacation days per year;
  • 10 National holidays;
  • Unlimited sick leave;
  • Private medical insurance compensation;
  • Co-working and gym/sports reimbursement;
  • Reward for the most innovative patentable idea.

Если просят выйти из iCloud, прислать код из SMS, запустить или установить что-то, перевести деньги — не соглашайтесь: это мошенничество.

Про зарплаты

Анонимные данные по зарплатам и грейдам.
Можно сверить вилку с рынком.

Посмотреть зарплаты

Если просят выйти из iCloud, прислать код из SMS, запустить или установить что-то, перевести деньги — не соглашайтесь: это мошенничество.