data engineer for unstructured data and AI pipelines
ориентир по рынку
вакансия
зп не указана
в среднем
205 915 ₽
мэтч
Загрузи резюме, чтобы видеть мэтчи с вакансией
подготовься к отклику
ai-инструменты
Чтобы адаптировать резюме под вакансию или составить сопроводительное письмо, загрузи резюме
О рекламодателе
ОБЩЕСТВО С ОГРАНИЧЕННОЙ ОТВЕТСТВЕННОСТЬЮ "ЦЕНТР НАЦИОНАЛЬНЫХ ИНТЕЛЛЕКТУАЛЬНЫХ СИСТЕМ" ИНН: 9704271170
описание
The project builds a private, access-scoped context layer over a European private investment group's data, with AI skills and agents built on top.
задачи
Set up call, email, Slack, and messenger ingestion with speaker attribution and reversible opt-out
Backfill historical email, Slack, board protocols, decks, and portfolio updates, ensuring they are parsed, deduplicated, and correctly dated
Build document parsing for PDFs, scanned board packs, spreadsheets, slide decks, and forwarded attachments
Implement identity and entity resolution across communication tools, calendars, portfolio companies, and CRM records
Build chunking and embedding pipelines and load vector and graph stores behind the architect-defined ontology
Implement incremental sync through the connector layer, handling edits and deletions without full re-crawls or silent drift
Attach access scope and provenance to every record during ingestion to support downstream permission-aware retrieval and audits
Run PII detection, redaction, and retention logic; provide the client's security function with evidence of what is stored, where, and for how long
Orchestrate monitored, repeatable pipelines with Airflow, Step Functions, or equivalent, and alert when a source stops flowing
Control cost and latency at volume through batching, incremental embedding, and storage tiering; report unit economics
Write runbooks so the client's team can operate the system after handover
требования
4+ Years in data engineering, including unstructured or semi-structured data work beyond warehouse modelling
Demonstrated experience integrating multiple third-party APIs into a coherent store, including historical backfill
Experience handling sensitive personal data in a regulated or security-sensitive environment
Comfortable working as the only data engineer in a small 2.5-FTE pod, at part-time allocation, without hand-holding
Strong Python and solid SQL
Experience with unstructured-data pipelines for transcripts, mail, chat, and documents, including parsing, normalisation, and deduplication
Experience with chunking strategies, vector stores such as pgvector, OpenSearch, or Pinecone-class systems, and graph-store loading
Experience integrating APIs and connectors at scale, including Google Workspace or M365, Slack, and CRM; rate limits, pagination, incremental cursors, and webhooks
Experience with entity resolution or record linkage, deterministic and fuzzy, without a clean shared key
Experience with Airflow, Step Functions, or equivalent orchestration and idempotent, restartable jobs
AWS and/or GCP data stack experience and comfort with private or VPC deployments
Practical knowledge of PII detection, redaction, encryption, and retention
Clear written English; able to prepare handover documents and work asynchronously in a small distributed pod
Knowledge of GDPR as applied to employee-generated data and EU data residency across multiple jurisdictions
Knowledge of data lineage, provenance, and audit patterns
Understanding of how retrieval quality depends on ingestion quality and sufficient RAG knowledge to make upstream choices
Будет плюсом: experience building pipelines feeding an LLM or retrieval system, Well-Architected security and cost practices, awareness of financial-services expectations
условия
Part-time engagement, 20 to 30 hours a week; allocation may flex above 0.5 FTE during Capture and Connect and settle back afterwards
Approximately eight to ten two-week sprints overall, with workload front-weighted to the first four or five