Чтобы адаптировать резюме под вакансию или составить сопроводительное письмо, загрузи резюме
описание
Intellias develops benchmark technological solutions and supports the digitalization of the world.
задачи
Provision and configure the Databricks environment in the customer’s cloud account with the customer’s platform team, including catalogs and permissions, compute, secrets, and experiment tracking
Build repeatable ingestion pipelines for raw verification outputs and document images, including parsing, pseudonymisation, data-quality checks, and schema handling
Implement feature tables for batch image-embedding jobs, exact and approximate linkage keys, cross-transaction and velocity aggregates, and pseudonymous entity resolution
Implement feature-selection and mixed-data clustering methods, run model comparisons and stability tests, and support density-model training runs
Orchestrate the end-to-end pipeline as a scheduled, idempotent, re-runnable workflow
Set up a data-quality and drift-monitoring prototype with alerting on feature and score tables
Implement access controls, provenance fields, lineage, and scripted deletion procedures; confirm that raw personal data does not reach the analytic layer
Write runbooks and README-level documentation, deliver a live walkthrough, and leave a single end-to-end Workflow that engineers can operate independently
требования
Experience with PySpark and Spark SQL, including nested-JSON flattening, pandas UDFs, window functions, partitioning, and performance tuning
Experience with Databricks Unity Catalog, S3 external locations, compute policies, Auto Loader, Delta, Workflows, Repos, and secret scopes
Experience with AWS S3, IAM roles or instance profiles, KMS decryption in jobs, and Secrets Manager; basic cost awareness
Knowledge of ML data engineering, including Medallion design, schema evolution, idempotent loads, data-quality profiling, and feature engineering in Unity Catalog
Experience implementing pseudonymisation in pipelines, including HMAC tokenization, normalization, phonetic keys, safe logging, and scripted deletion
Experience implementing entity resolution, including candidate generation, approximate keys, edit-distance tolerance, match scores, surrogate IDs, and conflict counting
Experience with GPU batch image processing using PyTorch or timm, including decryption, cropping, embeddings, PCA, and LSH bucketing
Hands-on experience with mixed-data clustering methods, including kmodes, gower, kmedoids, kamila, StepMix, and bootstrap ARI
Knowledge of feature selection methods, including mutual information, correlation pruning, PCA, and BPSO or GA search with mealpy
Experience with MLOps and monitoring using MLflow, Lakehouse Monitoring, SQL alerts, PSI, and JS divergence
Ability to write runbooks and README-level documentation
Будет плюсом: Experience with AssureID or identity-verification JSON outputs, PyOD, basic PyTorch training, and supporting work on VAEs