Если вы раньше входили через Google, сбросьте пароль для своей Gmail-почты через кнопку «Забыли пароль?» на экране входа. Затем войдите по email и новому паролю.
Если аккаунта ещё нет, зарегистрируйтесь с Gmail-почтой — после подтверждения почты мы предложим задать пароль.
Что нового
Загружаю обновления...
Что нового
Загружаю обновления...
Работа найдется быстрее с подпискойКандидат найдётся быстрее с подпиской
site reliability engineer for high-traffic applications
ориентир по рынку
вакансия
зп не указана
в среднем
325 071 ₽
мэтч
Загрузи резюме, чтобы видеть мэтчи с вакансией
подготовьтесь к отклику
ai-инструменты
Чтобы адаптировать резюме под вакансию или составить сопроводительное письмо, загрузите резюме
описание
Chess.com is a gaming, technology, and content company operating a global platform for playing, learning, and enjoying chess. It serves more than 250 million chess players worldwide through its products, content, and tools.
задачи
Design and implement multi-regional resilient infrastructure capable of handling millions of concurrent sessions and daily transactions across global data centers;
Lead the hybrid cloud migration strategy by integrating bare-metal data center resources with cloud services for performance and cost efficiency;
Own the on-call rotation and incident response procedures to ensure rapid resolution of critical system issues and maintain high-availability SLAs;
Architect monitoring and alerting systems to identify and resolve performance bottlenecks proactively;
Collaborate with development teams to implement infrastructure-as-code practices and deployment pipelines supporting continuous integration and delivery;
Optimize system performance through capacity planning, load testing, and resource allocation across distributed computing environments;
Establish and maintain security protocols and risk assessment procedures for infrastructure components and data protection;
Partner with engineering teams to design scalable solutions for high-traffic applications and real-time processing requirements;
Drive automation initiatives to reduce manual operational overhead and improve system reliability through scripting and configuration management;
Mentor team members on SRE best practices and contribute to infrastructure standards and documentation.
требования
Bachelor's degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience;
5+ Years of experience in site reliability engineering, DevOps, or infrastructure engineering roles;
Experience managing bare-metal server infrastructure and data center operations;
Strong proficiency with UNIX/Linux operating systems and command-line administration;
Experience with cloud platforms such as GCP, AWS, or Azure and infrastructure-as-code tools such as Terraform or CloudFormation;
Hands-on experience with configuration management systems such as Ansible, Chef, or Puppet;
Solid understanding of networking fundamentals, TCP/IP, HTTP/HTTPS, DNS, and network troubleshooting;
Experience with containerization and orchestration technologies such as Docker or Kubernetes;
Proficiency with monitoring and observability tools such as Datadog, Prometheus, Grafana, or ELK stack;
Experience with relational and NoSQL databases, including performance optimization and scaling strategies;
Strong collaboration and communication skills for working effectively in a distributed team environment;
Demonstrated ownership and accountability for system reliability and performance;
Nice to have: Advanced knowledge of CDNs and edge computing, server-side automation and scripting with Python, Go, or Bash, high-availability architectures and disaster recovery planning, security frameworks and compliance requirements, game server infrastructure or real-time application hosting, database administration and optimization for high-concurrency applications, CI/CD pipelines and deployment automation, capacity planning and performance testing tools, experience in a fully remote distributed work environment, continuous learning mindset with interest in emerging infrastructure technologies.