site reliability engineer for healthcare voice services
ориентир по рынку
вакансия
зп не указана
в среднем
328 556 ₽
мэтч
Загрузи резюме, чтобы видеть мэтчи с вакансией
подготовься к отклику
ai-инструменты
Чтобы адаптировать резюме под вакансию или составить сопроводительное письмо, загрузи резюме
описание
The platform supports medical voice services and workflow automation for healthcare systems in the US.
задачи
Define service-level objectives for production services
Build and maintain reliability dashboards
Instrument metrics, logs, and traces
Monitor workflow, voice, and third-party service dependencies
Lead incident response during agreed coverage hours
Write post-incident reviews
Improve deployment safety with rollback and staged releases
Maintain change records and release processes
Test capacity, backups, and disaster recovery
Coordinate handover and escalation with engineering teams
Provide uptime and reliability evidence
Support continuous improvement of platform resilience
требования
Strong production experience in SRE, DevOps, or Backend Operations
Hands-on experience with monitoring, logging, alerting, and tracing
Experience with incident response or incident management
Clear responsibility for uptime, SLA/SLO, or service reliability
Experience defining or working with SLOs and production dashboards
Strong cloud infrastructure background
Experience with containers and deployment automation
Experience with Infrastructure as Code (IaC)
Strong CI/CD knowledge
Python or comparable scripting skills for operational troubleshooting and automation
Experience diagnosing and resolving production failures
Understanding of backup, recovery, and capacity testing
Ability to communicate clearly during incidents and provide concise written updates
The CV must clearly show direct involvement in incident response and accountability for uptime/SLA/SLO, not only infrastructure configuration
Candidates should be comfortable with regular communication and incident-related coverage during US/Arizona business hours
Будет плюсом: healthcare or another regulated production environment, WebRTC/real-time media/voice infrastructure, LLM-backed services, handling third-party provider outages, Grafana/Prometheus/Datadog or similar observability tools, customer-facing incident communication
условия
25 Calendar days of vacation and 5 additional paid sick days
Medical insurance
Corporate English courses
Corporate events and team-building activities
Support with professional certifications
Reimbursement for professional courses and training
Long-term international projects
Opportunities for professional and technical growth
Calls are expected during Arizona time (MST, UTC-7)