Если вы раньше входили через Google, сбросьте пароль для своей Gmail-почты через кнопку «Забыли пароль?» на экране входа. Затем войдите по email и новому паролю.
Если аккаунта ещё нет, зарегистрируйтесь с Gmail-почтой, после подтверждения почты мы предложим задать пароль.
Что нового
Загружаю обновления...
Что нового
Загружаю обновления...
Работа найдется быстрее с подпискойКандидат найдётся быстрее с подпиской
Чтобы адаптировать резюме под вакансию или составить сопроводительное письмо, загрузите резюме
описание
Visa sponsorship and relocation assistance are available from any location
Scopely is a global video game and interactive entertainment company that creates, develops, publishes, and live-operates games across mobile, web, PC, and console platforms. Its portfolio includes MONOPOLY GO!, Pokémon GO, Stumble Guys, and other games reaching hundreds of millions of players worldwide.
задачи
Operate, monitor, and continuously improve messaging infrastructure, currently focused on NATS cluster and NATS JetStream;
Design and own the observability layer for distributed backend systems, defining metrics, traces, and logs and driving their adoption;
Correlate infrastructure signals with application behaviour to diagnose root causes and guide teams toward fixes;
Define, implement, and maintain SLO-related frameworks for backend services and messaging pipelines;
Build and maintain reliability tooling, runbooks, and operational frameworks as internal products;
Partner with backend, infrastructure, and product engineers to shape reliability standards, share operational context, and influence architecture decisions;
Lead or contribute to incident response across messaging and backend layers;
Drive postmortems toward systemic fixes and embed preventative improvements into engineering workflows;
Navigate infrastructure and backend codebases to identify instrumentation gaps, inadequate infrastructure architecture, missing error handling, and operational risks;
Contribute improvements in close collaboration with engineering teams.
требования
Strong background in Site Reliability Engineering, production operations, or backend engineering with a significant operational focus;
Hands-on production experience with NATS, NATS JetStream, or equivalent distributed messaging systems such as Apache Kafka or AWS Kinesis;
Experience with Leader Election, RAFT consensus, and log replication;
Ability to reason about application codebases such as C#, Go, Python, or similar;
Ability to identify instrumentation gaps, operational anti-patterns, and code-level root causes;
Strong experience designing and implementing metrics, logs, and traces strategies across distributed systems;
Experience debugging complex distributed systems across infrastructure and application layers;
Solid understanding of cloud infrastructure, preferably AWS, and containerized workloads such as ECS or Kubernetes/EKS;
Strong communication skills and ability to translate between infrastructure and application engineers and explain operational risk to non-technical stakeholders;
Infrastructure-as-Code-first approach to managing infrastructure, SLOs, monitors, and dashboards through code;
Nice to have: Direct NATS JetStream experience, Datadog or equivalent, database-level operational concerns, data modeling, mentoring engineers, driving reliability culture across teams.
условия
Remote or hybrid work is available in Spain, Ireland, Portugal, or the UK;
Visa sponsorship and relocation assistance are available from any location.