Чтобы адаптировать резюме под вакансию или составить сопроводительное письмо, загрузи резюме
описание
The client is a global consulting and technology organisation building an entity resolution platform that uses AI, machine learning, search technologies, and distributed systems to process hundreds of millions of records.
задачи
Design and operate large-scale data processing systems using Python, Spark, and SQL
Develop entity matching, deduplication, clustering, and search capabilities
Build and optimise data pipelines using Spark, Airflow, and DBT
Improve matching accuracy, recall, performance, and scalability
Develop analytical data models and optimise large-scale SQL workloads
Implement validation frameworks and monitor data quality metrics
требования
5+ Years of software engineering experience
Strong Python development skills
Hands-on production experience with Apache Spark / Databricks
Strong SQL skills and experience with large-scale datasets
Solid understanding of distributed systems and data-intensive applications
Experience with similarity matching, search, entity resolution, or related domains
Ability to balance accuracy, performance, scalability, and cost
Будет плюсом: Snowflake or similar cloud data warehouse platforms, Airflow, DBT or equivalent orchestration tools, entity resolution, record linkage or deduplication systems, embeddings, vector search, semantic matching or ML-assisted pipelines, Elasticsearch, OpenSearch or vector databases, Azure, AWS, Kubernetes, serverless technologies, GitHub Actions