Если вы раньше входили через Google, сбросьте пароль для своей Gmail-почты через кнопку «Забыли пароль?» на экране входа. Затем войдите по email и новому паролю.
Если аккаунта ещё нет, зарегистрируйтесь с Gmail-почтой, после подтверждения почты мы предложим задать пароль.
Что нового
Загружаю обновления...
Что нового
Загружаю обновления...
Работа найдется быстрее с подпискойКандидат найдётся быстрее с подпиской
Чтобы адаптировать резюме под вакансию или составить сопроводительное письмо, загрузите резюме
описание
A global leader in AI video generation, trusted by over 90% of Fortune 100 companies to transform how teams communicate and create content, operates as a high-growth Series E unicorn valued at $4B+ with over $500M raised from premier investors including Accel, Kleiner Perkins, and Nvidia's VC arm, and pushes the boundaries of generative AI.
задачи
Design and improve platform systems for model training, evaluation, and production serving;
Build robust infrastructure and tooling to make ML workloads scalable, reliable, and cost-efficient;
Architect the deployment and serving of ML models across research and production environments;
Improve scheduling, monitoring, and debugging for GPU and cloud-based workloads;
Develop internal abstractions, developer tools, and agentic systems to reduce operational overhead;
Drive continuous improvements across observability, automation, reliability, and developer experience (DX);
Collaborate closely with ML researchers and product engineers to turn pain points into robust platform capabilities;
Contribute to technical direction and make pragmatic architectural trade-offs.
требования
Strong experience building and operating complex, high-load production systems;
Deep systems mindset: ability to analyze bottlenecks, failure modes, and resource usage;
Solid hands-on experience with Linux, cloud infrastructure, and infrastructure automation;
Extensive experience with Kubernetes (K8s) and operating distributed workloads in production;
Strong coding skills in Python (or similar) for backend systems and tooling;
Proven experience building internal platforms, infrastructure abstractions, or developer tools;
Pragmatic approach to problem-solving with a focus on reliability without over-engineering;
Strong ownership and comfort working in ambiguous environments;
Upper-intermediate English (B2+);
Nice to have: direct experience operating ML infrastructure, GPU clusters, or model serving systems in production, familiarity with workflow orchestration systems (e.g., Temporal), experience building LLM-powered or agentic internal tools, strong background in observability and debugging distributed systems (Datadog, Prometheus, etc.), hands-on experience with Terraform, GitHub Actions, and CI/CD pipelines, experience bridging the gap between research and production engineering.