Чтобы адаптировать резюме под вакансию или составить сопроводительное письмо, загрузи резюме
О рекламодателе
ОБЩЕСТВО С ОГРАНИЧЕННОЙ ОТВЕТСТВЕННОСТЬЮ "ЦЕНТР НАЦИОНАЛЬНЫХ ИНТЕЛЛЕКТУАЛЬНЫХ СИСТЕМ" ИНН: 9704271170
описание
The organisation conducts AI research and engineering and builds data infrastructure to support next-generation AI models at massive scale.
задачи
Design, build, and scale distributed data-processing infrastructure for petabyte-scale workloads
Develop execution layers for pipeline stages with different computational requirements, selecting appropriate processing engines and technologies
Optimise compute utilisation across large-scale workloads, including GPU-accelerated data-processing stages
Build reliable pipelines with failure recovery, restartability, monitoring, and observability
Profile large-scale workloads to identify throughput bottlenecks and optimise compute utilisation and infrastructure costs
Establish systems for dataset versioning, lineage, reproducibility, and traceability across the data-processing lifecycle
Ensure intermediate and final datasets can be reproduced against specific research requirements
Build and maintain data infrastructure integrated with large-scale machine learning training pipelines
Evaluate technologies and architectural approaches to improve scalability, reliability, performance, and cost efficiency
Establish technical documentation, standards, and best practices across the data platform
требования
Bachelor's degree or higher in Computer Science, Computer Engineering, Software Engineering, or a related discipline
At least 5+ years of full-time experience designing, implementing, and operating large-scale distributed data-processing systems or ML data infrastructure
Demonstrated ownership of production data-processing systems under significant throughput, reliability, and cost pressures
Strong experience with distributed processing frameworks such as Ray, Apache Spark, or Flink
Experience with workflow orchestration and large-scale pipeline execution
Strong understanding of performance profiling, throughput optimisation, and distributed computing
Experience with Apache Arrow, Parquet, or comparable high-performance data formats
Experience with dataset versioning, lineage, reproducibility, and traceability
Experience with workload managers or schedulers such as Ray, Kubernetes, or Slurm
Familiarity with containerisation technologies including Docker or Enroot
Experience designing systems for fault tolerance, observability, and failure recovery
Strong software engineering skills and ability to make sustainable technical decisions in rapidly evolving environments
Будет плюсом: Experience supporting GPU-accelerated data-processing workloads, experience with cloud infrastructure and large-scale compute environments, familiarity with vector databases and modern data infrastructure, experience with infrastructure-as-code, experience working directly with ML training infrastructure or frontier AI workloads