вчера

data engineer for unstructured data and AI pipelines

ориентир по рынку
вакансия зп не указана
в среднем 205 915 ₽
Загрузи резюме, чтобы видеть мэтчи с вакансией

подготовься к отклику

ai-инструменты

Чтобы адаптировать резюме под вакансию или составить сопроводительное письмо, загрузи резюме

Рекламный баннер: ОБЩЕСТВО С ОГРАНИЧЕННОЙ ОТВЕТСТВЕННОСТЬЮ "ЦЕНТР НАЦИОНАЛЬНЫХ ИНТЕЛЛЕКТУАЛЬНЫХ СИСТЕМ"
О рекламодателе
ОБЩЕСТВО С ОГРАНИЧЕННОЙ ОТВЕТСТВЕННОСТЬЮ "ЦЕНТР НАЦИОНАЛЬНЫХ ИНТЕЛЛЕКТУАЛЬНЫХ СИСТЕМ"
ИНН: 9704271170

описание

The project builds a private, access-scoped context layer over a European private investment group's data, with AI skills and agents built on top.

задачи

  • Set up call, email, Slack, and messenger ingestion with speaker attribution and reversible opt-out
  • Backfill historical email, Slack, board protocols, decks, and portfolio updates, ensuring they are parsed, deduplicated, and correctly dated
  • Build document parsing for PDFs, scanned board packs, spreadsheets, slide decks, and forwarded attachments
  • Implement identity and entity resolution across communication tools, calendars, portfolio companies, and CRM records
  • Build chunking and embedding pipelines and load vector and graph stores behind the architect-defined ontology
  • Implement incremental sync through the connector layer, handling edits and deletions without full re-crawls or silent drift
  • Attach access scope and provenance to every record during ingestion to support downstream permission-aware retrieval and audits
  • Run PII detection, redaction, and retention logic; provide the client's security function with evidence of what is stored, where, and for how long
  • Orchestrate monitored, repeatable pipelines with Airflow, Step Functions, or equivalent, and alert when a source stops flowing
  • Control cost and latency at volume through batching, incremental embedding, and storage tiering; report unit economics
  • Write runbooks so the client's team can operate the system after handover

требования

  • 4+ Years in data engineering, including unstructured or semi-structured data work beyond warehouse modelling
  • Demonstrated experience integrating multiple third-party APIs into a coherent store, including historical backfill
  • Experience handling sensitive personal data in a regulated or security-sensitive environment
  • Comfortable working as the only data engineer in a small 2.5-FTE pod, at part-time allocation, without hand-holding
  • Strong Python and solid SQL
  • Experience with unstructured-data pipelines for transcripts, mail, chat, and documents, including parsing, normalisation, and deduplication
  • Experience with chunking strategies, vector stores such as pgvector, OpenSearch, or Pinecone-class systems, and graph-store loading
  • Experience integrating APIs and connectors at scale, including Google Workspace or M365, Slack, and CRM; rate limits, pagination, incremental cursors, and webhooks
  • Experience with entity resolution or record linkage, deterministic and fuzzy, without a clean shared key
  • Experience with Airflow, Step Functions, or equivalent orchestration and idempotent, restartable jobs
  • AWS and/or GCP data stack experience and comfort with private or VPC deployments
  • Practical knowledge of PII detection, redaction, encryption, and retention
  • Clear written English; able to prepare handover documents and work asynchronously in a small distributed pod
  • Knowledge of GDPR as applied to employee-generated data and EU data residency across multiple jurisdictions
  • Knowledge of data lineage, provenance, and audit patterns
  • Understanding of how retrieval quality depends on ingestion quality and sufficient RAG knowledge to make upstream choices
  • Будет плюсом: experience building pipelines feeding an LLM or retrieval system, Well-Architected security and cost practices, awareness of financial-services expectations

условия

  • Part-time engagement, 20 to 30 hours a week; allocation may flex above 0.5 FTE during Capture and Connect and settle back afterwards
  • Approximately eight to ten two-week sprints overall, with workload front-weighted to the first four or five
  • Active project; start ASAP

Если просят выйти из iCloud, прислать код из SMS, запустить или установить что-то, перевести деньги — не соглашайся: это мошенничество.

Про зарплаты

Анонимные данные по зарплатам и грейдам.
Можно сверить вилку с рынком.

Посмотреть зарплаты

Если просят выйти из iCloud, прислать код из SMS, запустить или установить что-то, перевести деньги — не соглашайся: это мошенничество.