Если вы раньше входили через Google, сбросьте пароль для своей Gmail-почты через кнопку «Забыли пароль?» на экране входа. Затем войдите по email и новому паролю.
Если аккаунта ещё нет, зарегистрируйтесь с Gmail-почтой, после подтверждения почты мы предложим задать пароль.
Что нового
Загружаю обновления...
Что нового
Загружаю обновления...
Работа найдется быстрее с подпискойКандидат найдётся быстрее с подпиской
Чтобы адаптировать резюме под вакансию или составить сопроводительное письмо, загрузите резюме
описание
No description
задачи
Design a data lake to process and store structured and unstructured data from raw sources;
Develop data models to store data and provide access for business purposes;
Develop integrations with various systems to provide a unified view of key indicators for decision-making;
Develop ETL pipelines to extract, transform, and load data;
Develop and monitor data pipelines to ensure data quality and integrity;
Automate and optimize data transformation processes;
Collaborate with other teams to ensure data is available and accessible for analytics and decision-making.
требования
Degree in computer science, information science, engineering, mathematics, or a related technical discipline;
5+ Years of experience with SQL and NoSQL technologies;
Strong experience with Oracle and PostgreSQL databases;
Extensive ETL development experience and strong Python programming skills;
Advanced knowledge of PySpark for large-scale distributed data processing;
Proficiency in Airflow for scheduling, orchestrating, and monitoring complex ETL/ELT workflows;
Expertise in Kafka for event streaming and messaging pipelines;
Experience with MinIO (S3-compatible storage), Parquet, and Iceberg;
Familiarity with Lakehouse concepts;
Deep understanding of MPP architecture concepts;
Nice to have: familiarity with Impala for interactive SQL querying on Big Data, experience with Greenplum, Docker and Kubernetes, CDC tools such as Debezium and Oracle GoldenGate, data visualization tools such as Tableau and PowerBI, DataOps and DevOps principles, data modeling and database design, data engineering best practices including data security, data access control and data governance, Grafana, GitLab, mentoring junior data engineers, and collaboration with Data Scientists, Analysts, and DevOps.