11 авг

site reliability engineer for financial analysis

ориентир по рынку
вакансия зп не указана
в среднем 351 041 ₽
Загрузи резюме, чтобы видеть мэтчи с вакансией

подготовьтесь к отклику

ai-инструменты

Чтобы адаптировать резюме под вакансию или составить сопроводительное письмо, загрузите резюме

описание

TradingView is a financial analysis and charting platform used by more than 100M users across 180+ countries. It provides advanced charting, market data, collaboration, publishing, and financial visualization tools for traders, investors, and companies.

задачи

  • Investigate production incidents and drive them through resolution until all consequences are fully addressed;
  • Perform root cause analysis and participate in postmortem reviews with engineering and product teams;
  • Track and drive corrective actions resulting from incidents and postmortem activities;
  • Develop and improve monitoring, alerting, and observability for assigned services;
  • Define and maintain SLI/SLOs, analyze reliability metrics, and monitor service health objectives;
  • Take ownership of SLA compliance for assigned frontend and backend services;
  • Define and implement service-level availability requirements with engineering and product owners;
  • Identify monitoring and observability gaps and implement improvements to reduce detection and recovery times;
  • Develop and maintain runbooks, troubleshooting guides, and service recovery procedures;
  • Participate in incident validation and escalation reviews;
  • Review and maintain operational and technical documentation;
  • Analyze service performance, resource utilization, and capacity-related risks;
  • Contribute to automation of diagnostics, incident response, and operational workflows;
  • Share operational knowledge and train duty engineers on new tools, procedures, and best practices;
  • Participate in on-call rotations and assist with complex or cross-service incidents.

требования

  • Experience as an SRE, Reliability Engineer, Production Engineer, Operations Engineer, or in a similar role;
  • Hands-on experience investigating production incidents and performing root cause analysis;
  • Strong understanding of monitoring, alerting, and observability principles;
  • Experience working with metrics, logs, and distributed tracing systems;
  • Knowledge of SLA, SLI, SLO, and error budget concepts;
  • Experience creating and maintaining operational documentation and runbooks;
  • Strong analytical and troubleshooting skills with the ability to drive investigations to actionable outcomes;
  • Understanding of distributed systems and high-load environments;
  • Experience automating operational and repetitive tasks;
  • Nice to have: CKA or CKAD certifications, experience with Prometheus, Grafana, OpenTelemetry, or similar observability platforms, experience implementing reliability practices within engineering organizations, experience with incident management and problem management processes, understanding of Google Four Golden Signals, experience using AI-assisted tools for incident investigation and diagnostics, experience operating large-scale distributed systems.

условия

  • Flexible working hours;
  • Relocation support and private health insurance;
  • Performance-based bonuses;
  • TradingView Premium access;
  • Learning, mentorship, and long-term career growth;
  • Regular team events and company-wide meetups.

Если просят выйти из iCloud, прислать код из SMS, запустить или установить что-то, перевести деньги — не соглашайтесь: это мошенничество.

Про зарплаты

Анонимные данные по зарплатам и грейдам.
Можно сверить вилку с рынком.

Посмотреть зарплаты

Если просят выйти из iCloud, прислать код из SMS, запустить или установить что-то, перевести деньги — не соглашайтесь: это мошенничество.