Чтобы адаптировать резюме под вакансию или составить сопроводительное письмо, загрузи резюме
описание
CloudLinux builds Linux infrastructure and security products for hosting companies and data centers. Its Infrastructure Department runs an observability platform, GitLab, CI runners, engineering services, and provisioning and configuration automation.
задачи
Run the observability platform, keep it healthy, onboard teams, monitor cost and capacity, and maintain alerting
Run GitLab and the CI runner fleet, including upgrades, capacity, access, backups, and restore drills
Keep other services healthy with production monitoring and runbooks
Deploy requested services from scratch by researching options, choosing designs, and ensuring they are managed as code, monitored, backed up, and documented
Handle developers’ requests related to access, onboarding, pipeline issues, exporters, and dashboards, and turn recurring requests into self-service
Respond to incidents, diagnose and mitigate impact, restore services safely, complete root-cause analyses and post-mortems, and implement prevention or detection improvements
Ship changes as code through reviewed merge requests, planning and checking each change
Write runbooks, onboarding guides, maintenance notices, and status updates for engineers outside the team
Work with AI agents by delegating collection and drafting, reviewing their output, and recording lessons for the team
требования
Senior-level experience in infrastructure, platform, or site reliability engineering, including responsibility for keeping at least one production service running
Linux systems administration and debugging on bare metal and virtual machines
Production Kubernetes experience delivered through GitOps, including personally performing cluster upgrades
Infrastructure as code experience using Ansible and Terraform or OpenTofu, with changes reviewed in merge requests
Production GitLab administration and GitLab CI experience, self-hosted or SaaS; equivalent depth with another CI system is acceptable
Working knowledge of Prometheus and Grafana, including running them for a team, writing alert rules and dashboards, and reading PromQL
Ability to write technical explanations for engineers outside the team, including runbooks, notices, and responses to requests
Strong communication and interpersonal skills; able to understand product-team needs, agree on scope, priority, and timing, push back politely, and keep stakeholders informed
Advanced use of AI engineering assistants such as Claude and Codex, including providing context, breaking down tasks, designing agent loops, and delegating scoped end-to-end execution with clear stop conditions; able to explain, debug, and test automation and verify generated commands, scripts, and conclusions before production use
Upper-intermediate or higher English
Not suited to candidates seeking ticket-queue operations, a pure cloud or Kubernetes role, or responsibility as a DBA, network engineer, or security engineer. Recurring requests are expected to become self-service; bare metal and virtual machines are a significant part of the infrastructure, and other teams run their own systems
Будет плюсом: SLO and burn-rate alert design with data-sized thresholds, Kata Containers, Firecracker or gVisor, S3-compatible object storage such as Ceph RGW, AWS cost work, self-hosted Sentry or Kafka-, ClickHouse- and Redis-backed applications kept running under load, Python or Go for exporters and small internal services
условия
Flexible working hours
24 Paid vacation days per year, 10 national holidays, and unlimited sick leave
Private medical insurance compensation
Co-working and gym/sports reimbursement
Education budget
Opportunity to receive a reward for an idea the company can patent