Чтобы адаптировать резюме под вакансию или составить сопроводительное письмо, загрузи резюме
описание
The client is a global management consultancy that partners with major organizations across industries to address complex business challenges through strategy, transformation, and performance improvement. Its European hub supports operations across the EMEA region.
задачи
Design and improve entity matching, clustering, and deduplication algorithms at scale
Implement distributed matching approaches, including blocking strategies, multi-pass matching pipelines, and nearest-neighbor and similarity-based methods
Apply and combine rule-based, statistical, and ML-assisted techniques, including embeddings where relevant
Optimize candidate generation and scoring to balance accuracy, recall, performance, and cost
Translate algorithmic ideas into scalable implementations using Spark and SQL transformations
Experiment with different approaches and iterate based on performance metrics and results
Design, build, and operate a large-scale analytic system processing hundreds of millions to billions of records
Implement and optimize Apache Spark pipelines for entity matching, deduplication, and clustering
Build and maintain complex workflows using Airflow and DBT
Ensure pipelines are fault-tolerant, observable, and cost-efficient in distributed environments
Develop and maintain analytical SQL models using Snowflake or a similar cloud data warehouse
Optimize large joins, aggregations, and window functions over very large datasets
Design data models for matching pipelines and downstream consumers
Build validation logic and metrics to measure match rate, precision, recall, and accuracy
Support continuous improvements to the matching engine through iterative releases
Debug and resolve data quality issues across heterogeneous and imperfect data sources
требования
5+ Years of software development using Python, ideally in a team lead capacity
Solid understanding of distributed systems and algorithms, including partitioning, shuffles, joins, and scalability trade-offs
Experience building and working with complex data pipelines or data systems
Strong SQL skills, ideally with Snowflake or similar analytical databases
Strong hands-on production experience with Apache Spark and Databricks (PySpark or Scala)
Experience with AI/ML-assisted systems, including embeddings, inference, and re-ranking
Experience with or strong interest in fuzzy and semantic matching techniques, such as Levenshtein/edit distance, token-based similarity, BM25 or other lexical ranking methods, vector embeddings and cosine similarity, and approximate nearest-neighbor or vector search concepts
Strong willingness to learn and apply advanced semantic matching techniques if not already experienced
Будет плюсом: experience with Airflow or equivalent orchestration tools, entity resolution/deduplication/record linkage systems, search or retrieval systems such as Elasticsearch, OpenSearch, or vector databases, data pipelines at very large scale (100M+ records), data quality frameworks or validation automation or QA at scale, Kubernetes or containerized functions such as Azure Container Apps or AWS Fargate, serverless functions, Azure or other cloud platforms, GitHub Actions