Если вы раньше входили через Google, сбросьте пароль для своей Gmail-почты через кнопку «Забыли пароль?» на экране входа. Затем войдите по email и новому паролю.
Если аккаунта ещё нет, зарегистрируйтесь с Gmail-почтой — после подтверждения почты мы предложим задать пароль.
Что нового
Загружаю обновления...
Что нового
Загружаю обновления...
Работа найдется быстрее с подпискойКандидат найдётся быстрее с подпиской
service availability engineer for production AI systems
ориентир по рынку
вакансия
зп не указана
в среднем
325 071 ₽
мэтч
Загрузи резюме, чтобы видеть мэтчи с вакансией
подготовьтесь к отклику
ai-инструменты
Чтобы адаптировать резюме под вакансию или составить сопроводительное письмо, загрузите резюме
описание
EPAM's Operational Intelligence practice applies SRE principles and cloud-native practices to the lifecycle of production AI/ML and LLM systems. It develops AI Reliability Engineering capabilities for monitoring, optimizing, and securing production AI applications and agentic systems.
задачи
Instrument production LLM, RAG, and agentic applications with AI telemetry based on OpenTelemetry and APM-native AI monitoring;
Implement distributed tracing across multi-model chains, agent workflows, and retrieval-augmented generation pipelines to profile systemic latency and failure points;
Define and measure AI-native SLIs and SLOs, including TTFT, throughput, error and refusal rates, cost per request, semantic drift, hallucination boundaries, and contextual accuracy;
Set up structured semantic logging and prompt/response monitoring for quality analysis;
Build and run evaluation loops for output quality and safety using golden sets, LLM-as-a-judge, and Ragas/DeepEval-style frameworks, integrating them into CI/CD and runtime;
Track token-based cloud spend, model API rate limits, and quota consumption; drive AI cost optimization;
Configure AI gateways for API load balancing, failover, and fallback models across multiple LLM providers;
Implement guardrails for prompt injection and jailbreak filtering, output compliance, bias, and safety constraints;
Design detection, triage, restore, and problem management workflows for AI incidents; integrate autonomous AI agents into RCA to parse logs, form hypotheses, and correlate state changes;
Support rollback, canary, and fail-safe patterns for model, prompt, and configuration releases; maintain reproducibility through versioning of data, code, prompts, and models;
Build practice accelerators, reference architectures, and internal enablement materials; support presales and client assessments.
требования
4+ Years in SRE, DevOps, platform, or observability engineering, including hands-on work with production AI/ML or LLM workloads;
Solid SRE fundamentals: Golden Signals, SLI/SLO definition, error budgets and burn rate, incident lifecycle, and ITIL basics;
Strong Python skills for instrumentation, automation, and evaluation tooling;
Practical experience with OpenTelemetry and at least one APM/observability platform: New Relic, Datadog, Grafana LGTM stack, Splunk, or Elastic;
Hands-on production experience with at least one cloud platform, with Azure preferred over AWS or GCP, and Kubernetes;
Working understanding of LLM application architecture, including prompts, embeddings and vector stores, RAG, and agent orchestration with LangChain, LangGraph, or equivalent;
MLOps awareness covering model lifecycle, model endpoints, containerization, deployment, and rollback patterns;
Infrastructure as Code experience with Terraform and CI/CD experience with Azure DevOps, GitLab CI, or GitHub Actions;
B2+ English for clear written and spoken technical communication in a client-facing role;
Nice to have: AI-specific observability and evaluation tooling, distributed inference serving at scale, AI security, Databricks or Azure AI Foundry, relevant certifications, FinOps for AI workloads, Data Reliability Engineering, mentoring or team lead experience.
условия
No 24/7 on-call rotation;
Funded certification and enablement tracks for Anthropic/Claude, Databricks, and AI & Data Observability;
Cross-client exposure and a direct path into presales and solution engineering.