Чтобы адаптировать резюме под вакансию или составить сопроводительное письмо, загрузи резюме
описание
Promofy is a fast-growing B2B SaaS startup that provides loyalty infrastructure for iGaming platforms. Its modular, real-time, event-driven backend serves thousands of players across multiple tenant environments.
задачи
Operate and maintain AWS production infrastructure and Kubernetes / EKS
Maintain and improve Terraform infrastructure, Karpenter, autoscaling, resource allocation, and cluster capacity
Maintain and improve CI/CD pipelines
Monitor production infrastructure and proactively identify problems
Troubleshoot and resolve production infrastructure incidents
Maintain and improve monitoring, logging, metrics, and alerting
Build internal infrastructure and operational automation using Python
Automate repetitive operational and maintenance tasks
Improve infrastructure reliability, scalability, and cost efficiency
Maintain infrastructure security, IAM, networking, and access controls
Perform root-cause analysis and implement permanent solutions for recurring issues
Maintain backup and recovery procedures
Manage the infrastructure, monitor the health and capacity, understand connectivity and resource requirements, and troubleshoot infrastructure-related issues for PostgreSQL / AWS Aurora, Redis, and Kafka / AWS MSK
Independently propose and implement infrastructure improvements and reduce unnecessary operational work
Use AI-assisted engineering tools where they improve productivity while maintaining understanding and ownership of production changes
Take ownership of production issues, respond quickly to critical incidents during agreed working hours, investigate infrastructure problems independently, and follow incidents through to resolution
Prevent recurring incidents rather than repeatedly fixing symptoms
требования
Hands-on experience with AWS, Kubernetes / EKS, Terraform, Karpenter, Python, CI/CD, Docker / containerized workloads, Linux, AWS networking and IAM, production monitoring, logging, alerting, and troubleshooting production environments
Production-grade AWS and Kubernetes experience is essential
Understand how production infrastructure components affect one another
Be proactive, automation-driven, reliable, innovative, and independent
Be available during agreed working hours and respond quickly to critical production incidents
Take ownership of production issues and follow critical incidents through to resolution