Чтобы адаптировать резюме под вакансию или составить сопроводительное письмо, загрузи резюме
О рекламодателе
ОБЩЕСТВО С ОГРАНИЧЕННОЙ ОТВЕТСТВЕННОСТЬЮ "ЦЕНТР НАЦИОНАЛЬНЫХ ИНТЕЛЛЕКТУАЛЬНЫХ СИСТЕМ" ИНН: 9704271170
описание
EPAM provides enterprise software products, open source solutions, and accelerators.
задачи
Collaborate with development, security, quality, and operations teams to implement SRE practices and ensure system reliability;
Define and support the required level of reliability, availability, and performance for services and applications;
Troubleshoot, mitigate, and support fixing infrastructure and application issues in a timely manner;
Implement monitoring systems for infrastructure and application reliability;
Define and track Service Level Objectives (SLOs);
Manage error budgets;
Reduce toil through automation;
Deploy, maintain, and automate infrastructure and application environments;
Enhance system reliability and optimize resource utilization;
Drive technology initiatives and maximize their impact across the organization.
требования
Bachelor’s degree in Computer Science, Engineering, or a related field;
Proven experience with any cloud platform: AWS, GCP, or Azure;
Experience implementing SRE practices, including SLO/SLI, error budgets, postmortems, reducing toil, capacity planning, and incident management;
Knowledge of Python or another scripting/programming language;
Strong background in monitoring tools;
Proficiency in CI/CD tools, infrastructure as code, and configuration management;
Solid knowledge of container orchestration technologies, including Kubernetes and Docker;
Nice to have: Expertise in deployment and management of LLMs, including RAG; certification in Kubernetes, AWS/GCP/Azure, or similar technologies; proven experience in DevOps; knowledge of managing and optimizing AI/ML models in production environments, including basic deployment, monitoring, and maintenance.