Если вы раньше входили через Google, сбросьте пароль для своей Gmail-почты через кнопку «Забыли пароль?» на экране входа. Затем войдите по email и новому паролю.
Если аккаунта ещё нет, зарегистрируйтесь с Gmail-почтой, после подтверждения почты мы предложим задать пароль.
Что нового
Загружаю обновления...
Что нового
Загружаю обновления...
Работа найдется быстрее с подпискойКандидат найдётся быстрее с подпиской
Чтобы адаптировать резюме под вакансию или составить сопроводительное письмо, загрузите резюме
описание
EPAM develops enterprise software products, open source solutions, and accelerators, including solutions for production AI systems.
задачи
Own end-to-end deployment, including infrastructure as code, CI/CD pipelines, and environment management on Azure;
Build LLM-aware observability with model-call traces, production quality signals, drift detection, guardrail-trigger rates, and cost and latency dashboards;
Define and defend SLOs for availability, latency, and quality objectives, balancing delivery speed against stability with data-driven error budgets;
Run incident management, including on-call models, pre-incident runbooks, and blameless postmortems;
Manage intelligence costs by monitoring token economics, configuring budgets and alerts, and conducting proactive capacity planning;
Maintain the security posture through patching, secret rotation, access reviews, and audit readiness across the solution lifecycle;
Shape operability requirements before handover and ensure their inclusion in the pod's definition of done;
Run hypercare jointly and sign off on what will be operated;
Feed operational patterns, failure modes, and cost learnings back to the pods and the Architect.
требования
5+ Years of experience operating cloud production systems, including scaling, defining SLOs, managing on-call rotations, and automating manual work;
Expertise in Azure IaaS/PaaS operations, infrastructure as code with Terraform/Bicep, and CI/CD tooling;
Proficiency with observability stacks, LLM tracing, container orchestration, and Python/Bash automation;
Knowledge of FinOps basics for AI workloads;
Familiarity with daily AI use in operations, including incident triage, runbook drafting, log analysis, and automation code.