NDA
28 сен

ml engineer for entity resolution

ориентир по рынку
вакансия зп не указана
в среднем 292 674 ₽
Загрузи резюме, чтобы видеть мэтчи с вакансией

подготовься к отклику

ai-инструменты

Чтобы адаптировать резюме под вакансию или составить сопроводительное письмо, загрузи резюме

описание

The client is a global management consultancy that partners with major organizations across industries to address complex business challenges through strategy, transformation, and performance improvement. Its European hub supports operations across the EMEA region.

задачи

  • Design and improve entity matching, clustering, and deduplication algorithms at scale
  • Implement distributed matching approaches, including blocking strategies, multi-pass matching pipelines, and nearest-neighbor and similarity-based methods
  • Apply and combine rule-based, statistical, and ML-assisted techniques, including embeddings where relevant
  • Optimize candidate generation and scoring to balance accuracy, recall, performance, and cost
  • Translate algorithmic ideas into scalable implementations using Spark and SQL transformations
  • Experiment with different approaches and iterate based on performance metrics and results
  • Design, build, and operate a large-scale analytic system processing hundreds of millions to billions of records
  • Implement and optimize Apache Spark pipelines for entity matching, deduplication, and clustering
  • Build and maintain complex workflows using Airflow and DBT
  • Ensure pipelines are fault-tolerant, observable, and cost-efficient in distributed environments
  • Develop and maintain analytical SQL models using Snowflake or a similar cloud data warehouse
  • Optimize large joins, aggregations, and window functions over very large datasets
  • Design data models for matching pipelines and downstream consumers
  • Build validation logic and metrics to measure match rate, precision, recall, and accuracy
  • Support continuous improvements to the matching engine through iterative releases
  • Debug and resolve data quality issues across heterogeneous and imperfect data sources

требования

  • 5+ Years of software development using Python, ideally in a team lead capacity
  • Solid understanding of distributed systems and algorithms, including partitioning, shuffles, joins, and scalability trade-offs
  • Experience building and working with complex data pipelines or data systems
  • Strong SQL skills, ideally with Snowflake or similar analytical databases
  • Strong hands-on production experience with Apache Spark and Databricks (PySpark or Scala)
  • Experience with AI/ML-assisted systems, including embeddings, inference, and re-ranking
  • Experience with or strong interest in fuzzy and semantic matching techniques, such as Levenshtein/edit distance, token-based similarity, BM25 or other lexical ranking methods, vector embeddings and cosine similarity, and approximate nearest-neighbor or vector search concepts
  • Strong willingness to learn and apply advanced semantic matching techniques if not already experienced
  • Будет плюсом: experience with Airflow or equivalent orchestration tools, entity resolution/deduplication/record linkage systems, search or retrieval systems such as Elasticsearch, OpenSearch, or vector databases, data pipelines at very large scale (100M+ records), data quality frameworks or validation automation or QA at scale, Kubernetes or containerized functions such as Azure Container Apps or AWS Fargate, serverless functions, Azure or other cloud platforms, GitHub Actions

условия

Условий нет

Если просят выйти из iCloud, прислать код из SMS, запустить или установить что-то, перевести деньги — не соглашайся: это мошенничество.

Спроси Хайрика про вакансию

Сверит с твоим резюме, подскажет вилку и вопросы на собесе.

Если просят выйти из iCloud, прислать код из SMS, запустить или установить что-то, перевести деньги — не соглашайся: это мошенничество.