site reliability engineer infrastructure reliability
ориентир по рынку
вакансия
зп не указана
в среднем
353 643 ₽
мэтч
Загрузи резюме, чтобы видеть мэтчи с вакансией
подготовься к отклику
ai-инструменты
Чтобы адаптировать резюме под вакансию или составить сопроводительное письмо, загрузи резюме
описание
No description
задачи
Collaborate with development, security, quality, and operations teams to implement SRE practices and ensure system reliability;
Define and support the required level of reliability, availability, and performance for services and applications;
Troubleshoot, mitigate, and support fixing infrastructure and application issues in a timely manner;
Implement monitoring systems for infrastructure and application reliability.
требования
Bachelor’s degree in Computer Science, Engineering, or a related field;
Proven experience with a cloud platform such as AWS, GCP, or Azure;
Experience implementing SRE practices, including SLO/SLI, error budgets, postmortems, toil reduction, capacity planning, and incident management;
Knowledge of Python or another scripting or programming language;
Strong background in monitoring tools;
Proficiency in CI/CD tools, infrastructure as code, and configuration management;
Solid knowledge of container orchestration technologies such as Kubernetes and Docker;
Nice to have: LLM deployment and management expertise, including RAG; Kubernetes, AWS/GCP/Azure, or similar certifications; DevOps experience; experience managing and optimizing AI/ML models in production, including basic deployment, monitoring, and maintenance.