mlops engineer
генерация резюме под вакансию
сопроводительное письмо
описание
A global leader in AI video generation, trusted by over 90% of Fortune 100 companies to transform how teams communicate and create content, operates as a high-growth Series E unicorn valued at $4B+ with over $500M raised from premier investors including Accel, Kleiner Perkins, and Nvidia's VC arm, and pushes the boundaries of generative AI.
задачи
- Design and improve platform systems for model training, evaluation, and production serving;
- Build robust infrastructure and tooling to make ML workloads scalable, reliable, and cost-efficient;
- Architect the deployment and serving of ML models across research and production environments;
- Improve scheduling, monitoring, and debugging for GPU and cloud-based workloads;
- Develop internal abstractions, developer tools, and agentic systems to reduce operational overhead;
- Drive continuous improvements across observability, automation, reliability, and developer experience (DX);
- Collaborate closely with ML researchers and product engineers to turn pain points into robust platform capabilities;
- Contribute to technical direction and make pragmatic architectural trade-offs.
требования
- Strong experience building and operating complex, high-load production systems;
- Deep systems mindset: ability to analyze bottlenecks, failure modes, and resource usage;
- Solid hands-on experience with Linux, cloud infrastructure, and infrastructure automation;
- Extensive experience with Kubernetes (K8s) and operating distributed workloads in production;
- Strong coding skills in Python (or similar) for backend systems and tooling;
- Proven experience building internal platforms, infrastructure abstractions, or developer tools;
- Pragmatic approach to problem-solving with a focus on reliability without over-engineering;
- Strong ownership and comfort working in ambiguous environments;
- Upper-intermediate English (B2+);
- Nice to have: direct experience operating ML infrastructure, GPU clusters, or model serving systems in production, familiarity with workflow orchestration systems (e.g., Temporal), experience building LLM-powered or agentic internal tools, strong background in observability and debugging distributed systems (Datadog, Prometheus, etc.), hands-on experience with Terraform, GitHub Actions, and CI/CD pipelines, experience bridging the gap between research and production engineering.
условия
- No conditions specified
навыки
Если просят войти через iCloud, отправить коды из SMS, запустить код, что-то установить, перевести деньги или сделать что угодно, связанное с деньгами, не соглашайтесь: это признаки мошенничества.