Чтобы адаптировать резюме под вакансию или составить сопроводительное письмо, загрузи резюме
О рекламодателе
ОБЩЕСТВО С ОГРАНИЧЕННОЙ ОТВЕТСТВЕННОСТЬЮ "ЦЕНТР НАЦИОНАЛЬНЫХ ИНТЕЛЛЕКТУАЛЬНЫХ СИСТЕМ" ИНН: 9704271170
описание
Nebius builds a full-stack AI cloud platform that supports developers and enterprises from data and model training through production deployment. Its platform provides AI/ML infrastructure across compute, storage, networking, and applied AI.
задачи
Own region and GPU cluster deployment projects from kickoff through production handover
Coordinate capacity acceptance, including infrastructure availability, quota allocations, monitoring and alerting, and documented sign-off before regions go to customers
Coordinate acceptance testing across compute, storage, networking, and GPU engineering teams, and clarify test ownership
Act as the real-time escalation point during live deployments, resolving blockers such as hardware failures, provisioning issues, and capacity conflicts with on-call teams
Define and maintain clear delivery contracts and responsibilities between teams
Drive continuous improvement and automation of repetitive deployment and acceptance-testing work
Run structured, high-signal status communications to keep stakeholders informed about deployment progress
требования
Proven experience leading technical projects or engineering teams, ideally in infrastructure, cloud, or datacenter environments
Ability to coordinate multiple engineering teams and drive alignment without formal authority
Calm, decisive communication under pressure, including during live incidents
Strong focus on clear ownership and documentation over ad hoc firefighting
Fluent English
Applicants must be authorized to work in the country where they apply and provide proof of employment eligibility as a condition of hire
Будет плюсом: experience with cloud/compute capacity management or GPU infrastructure, background in large-scale datacenter rollouts