Если вы раньше входили через Google, сбросьте пароль для своей Gmail-почты через кнопку «Забыли пароль?» на экране входа. Затем войдите по email и новому паролю.
Если аккаунта ещё нет, зарегистрируйтесь с Gmail-почтой — после подтверждения почты мы предложим задать пароль.
Что нового
Загружаю обновления...
Что нового
Загружаю обновления...
Работа найдется быстрее с подпискойКандидат найдётся быстрее с подпиской
Чтобы адаптировать резюме под вакансию или составить сопроводительное письмо, загрузите резюме
описание
EPAM delivers enterprise software products, open source solutions, and accelerators. The client-facing delivery team is modernizing a data platform by building cloud-native data pipelines on AWS and delivering curated data into Snowflake.
задачи
Design, build, and optimize scalable batch and incremental ETL/ELT pipelines using PySpark on AWS Glue;
Configure Glue jobs, crawlers, triggers, connections, bookmarks, workflows, and the Glue Data Catalog;
Tune workers, partitioning, and shuffle behavior for cost and performance optimization;
Model and load curated datasets into Snowflake with staging, transformation, and publishing layers;
Implement automated data quality and validation frameworks, including schema/contract enforcement and null, uniqueness, and referential checks;
Develop row-count and financial reconciliation processes, anomaly detection, and quarantine/reject handling;
Configure and extend Glue Data Quality DQDL rules according to requirements;
Write clean, modular, testable Python with unit and integration tests and reusable libraries;
Integrate pipelines with AWS services such as S3, IAM, Lambda, Athena, CloudWatch, Step Functions, and Secrets Manager;
Instrument observability through logging, metrics, alerting, and pipeline SLA monitoring;
Participate in code reviews, CI/CD automation, and documentation;
Engage directly with client stakeholders in requirements refinement, design walkthroughs, status reporting, and technical advisory within the workstream.
требования
3+ Years of production-level Python experience in data engineering, including OOP and functional patterns;
Expert knowledge of PySpark for distributed data processing and the DataFrame API;
Advanced proficiency in Snowflake, including data warehousing and staging/transformation layers;
Experience with AWS Glue, including job configuration, crawlers, Data Catalog, and DQDL;
Background in data quality engineering, including validation frameworks and reconciliation;
Proficiency in AWS services including S3, IAM, Lambda, Athena, and CloudWatch;
English proficiency at B2 level or higher;
Nice to have: Generative AI / LLM concepts, Airflow / Step Functions orchestration, Great Expectations or similar data quality frameworks, Terraform / CloudFormation.