Если вы раньше входили через Google, сбросьте пароль для своей Gmail-почты через кнопку «Забыли пароль?» на экране входа. Затем войдите по email и новому паролю.
Если аккаунта ещё нет, зарегистрируйтесь с Gmail-почтой, после подтверждения почты мы предложим задать пароль.
Что нового
Загружаю обновления...
Что нового
Загружаю обновления...
Работа найдется быстрее с подпискойКандидат найдётся быстрее с подпиской
Чтобы адаптировать резюме под вакансию или составить сопроводительное письмо, загрузите резюме
описание
EPAM develops enterprise software products, open source solutions, and accelerators. The role supports AI/GenAI applications for Risk-Based Quality Management in clinical trials.
задачи
Design and build RAG document ingestion pipelines for clinical trial quality data;
Build and manage vector databases using AWS OpenSearch for RAG-powered AI workflows;
Develop batch and streaming ETL/ELT pipelines from scratch for unstructured clinical data;
Build and expose data APIs for AI application consumption;
Optimize chunking strategies, embedding generation, and retrieval performance for RAG architectures;
Manage data quality, lineage, and governance for AI/ML data pipelines;
Deploy and maintain AWS data infrastructure;
Collaborate with Data Scientists and Backend Developers in an integrated pod team.
требования
5+ Years of hands-on data engineering experience at scale;
Expertise in RAG document ingestion pipelines, including chunking, embedding, and vector indexing;
Proficiency in AWS OpenSearch as a vector database for RAG workflows;
Advanced proficiency in Python, including SQL and Spark SQL;
Experience transforming unstructured data such as PDF and DOCX for RAG/LLM applications;
Familiarity with AWS Services including S3, Lambda, Glue, Athena, Bedrock, Step Functions, API Gateway, CloudWatch, and DynamoDB;
Knowledge of containerization with Docker;
Ability to build custom pipelines from scratch beyond configuring out-of-the-box services;
English proficiency at B2+ level;
Nice to have: pharmaceutical or life sciences background, Snowflake, Pinecone, SageMaker processing jobs, CI/CD tools such as Jenkins and Git/Bitbucket, infrastructure tools such as CDK or Terraform, clinical data standards including CDISC, ADaM, and SDTM.