Чтобы адаптировать резюме под вакансию или составить сопроводительное письмо, загрузи резюме
описание
Aristek Systems develops IT solutions for the US and European markets. Its project client provides technology-enabled solutions and services to higher education and nonprofit organizations, helping them optimize student enrollment and fundraising.
задачи
Design, develop, and maintain data pipelines and ETL/ELT processes for structured and unstructured data, ensuring data quality and integrity
Create and manage scalable data architectures for frequent transformations and manipulations across complex, multi-source environments
Collaborate with data scientists, analysts, and stakeholders to ensure pipelines meet business requirements
Monitor, troubleshoot, and optimize pipelines for large-scale datasets, including real-time and batch processing
Ensure compliance with data governance, security, and privacy standards, including data lineage, access control, and auditability
Stay current with data engineering technologies, AI frameworks, and best practices to improve infrastructure and workflows
требования
At least 2 years of experience
Experience building scalable, automated data pipelines using ETL/ELT methodologies and handling complex transformations and frequent data manipulations
Good experience with Databricks
Experience with relational databases, NoSQL, APIs, streaming data such as Apache Kafka, and unstructured data formats
Strong Python (pySpark, SparkSQL) and SQL skills
Hands-on experience with Apache Spark and cloud platforms such as IBM Cloud, AWS, Azure, or Google Cloud
Knowledge of Apache Airflow or Dagster for workflow orchestration, scheduling, dependency management, and failure recovery
Strong focus on data quality, data lineage, GDPR and CCPA compliance, and security best practices, including encryption and role-based access control
Ability to resolve challenges in large-scale, heterogeneous data ecosystems and ensure pipeline reliability and performance