platform engineer
генерация резюме под вакансию
сопроводительное письмо
описание
The company operates in the financial services sector and is looking for an AI Platforms Engineer to design, build, and operate the foundational platform that enables teams to develop, deploy, govern, and scale AI solutions reliably and securely.
задачи
- Design and build scalable AI platform capabilities for training, fine-tuning, inference, evaluation, and experimentation;
- Develop and operate shared platform services for model serving, vector databases, feature and data access, prompt and agent workflows, GPU workload orchestration, secrets, and configuration management;
- Build reusable MLOps and LLMOps pipelines for model packaging, deployment, rollback, versioning, and lifecycle management;
- Enable secure deployment and operation of open-source models, commercial model APIs, retrieval-augmented generation systems, and agent-based workloads;
- Create internal self-service tooling and templates for AI application teams;
- Implement platform controls for authentication and authorization, rate limiting and quota management, audit logging, data protection, policy enforcement, and guardrails;
- Build observability for AI workloads, including latency, throughput, token usage, GPU utilization, model and system health, drift, and quality indicators;
- Improve reliability and efficiency of AI infrastructure through automation, SRE practices, and performance tuning;
- Partner with data scientists, software engineers, architects, security teams, and business stakeholders to translate AI use cases into robust platform capabilities;
- Define standards and best practices for AI platform architecture, CI/CD, monitoring, governance, and operations;
- Support evaluation and integration of emerging AI infrastructure technologies, frameworks, and tools.
требования
- Strong experience with Kubernetes and containerized workloads in production;
- 5+ Years experience in production infrastructure as a Platform, SRE, DevOps, or MLOps engineer;
- Strong scripting and automation in Go and/or Python, including writing Kubernetes controllers or non-trivial operational tools;
- Experience with CI/CD pipelines, Infrastructure-as-Code, and GitOps such as GitLab CI, Terraform, Ansible, Helm, ArgoCD, or Flux;
- Solid understanding of Linux, networking, storage, and security;
- Experience with monitoring and observability tools such as Prometheus, Grafana, and OpenTelemetry;
- Understanding of the ML lifecycle including training, deployment, inference, evaluation, and monitoring;
- Experience building or operating shared platforms used by multiple teams;
- Ability to work closely with data scientists, ML engineers, software engineers, and security teams;
- Depth in at least one area including GPU infrastructure and inference ops, platform SRE at scale, MLOps, or API gateway operations;
- Nice to have: experience with LLM and GenAI platforms, experience with model serving tools such as vLLM, Triton, TGI, Ray, KServe, or TorchServe, familiarity with RAG, embeddings, reranking, and vector databases, experience with GPU infrastructure and scheduling AI workloads, experience integrating open-source models and commercial AI APIs, experience with identity and access management such as LDAP, SAML, OIDC, OAuth2.
условия
- No conditions specified
навыки
Если просят войти через iCloud, отправить коды из SMS, запустить код, что-то установить, перевести деньги или сделать что угодно, связанное с деньгами, не соглашайтесь: это признаки мошенничества.