сегодня

ml engineer for document intelligence

ориентир по рынку
вакансия зп не указана
в среднем 306 413 ₽
Загрузи резюме, чтобы видеть мэтчи с вакансией

подготовьтесь к отклику

ai-инструменты

Чтобы адаптировать резюме под вакансию или составить сопроводительное письмо, загрузите резюме

описание

The client is a global investment management company headquartered in London. It manages over $228 billion in assets and serves institutional investors, pension funds, wealth managers, and other sophisticated clients worldwide. It specializes in quantitative investing, alternative investments, systematic trading strategies, and technology-driven asset management, with data science, machine learning, and AI central to its investment and research processes. The project builds data foundations that make AI useful and safe inside regulated financial firms.

задачи

  • Build a versioned automated evaluation test suite from the client’s question bank and diagnostic results, establish a baseline, and use historical support records as a second ground-truth set
  • Improve the existing RAG service with hybrid retrieval, reranking, structure-aware chunks, context notes, early metadata filters, and reliability monitoring; measure every change against the suite before release
  • Build a metadata and entity extraction pipeline under a governed three-tier schema, with a stratified pilot, confidence-routed human review, and coverage dashboards
  • Build the synchronisation plane over the client’s content platform, including change notifications, delta polling, and scheduled full reviews, shared by retrieval, extraction, and downstream views
  • Build document lineage and temporal views, including supersession and amendment chains extracted only when stated in the text, a bitemporal effective-terms view, derived document status, and a queryable obligations register; provide human verification for high-stakes chains
  • Contribute to the controlled knowledge graph and query orchestration, including typed nodes and edges with provenance and confidence, routing between structured lookup, filtered retrieval, and lineage views, explicit completeness statements, and typed gaps reported as answers
  • Operate the delivered capabilities, gate releases on evaluation results, enforce freshness bounds by withholding stale data, maintain coverage and quality dashboards, and report corpus health checks to document owners

требования

  • 6+ Years of production Python development, including 2+ years of production LLM and RAG engineering
  • Strong experience building evaluation harnesses with versioned test suites derived from real question banks, separate retrieval and answer scoring, “not found” scoring, and delivery release gates; tools may include Langfuse, RAGAS, DeepEval, or similar, plus custom metrics
  • Experience with hybrid retrieval engineering, including keyword and semantic search, result merging, cross-encoder reranking, measured baseline tuning, and working with a platform-fixed embedding model and index
  • Experience with structure-aware document processing and chunking that keeps tables intact, including OCR for scanned documents and multilingual content
  • Experience with large-scale, schema-driven LLM extraction of attributes, entities, clauses, and relationships, including per-field confidence, calibrated thresholds, human review, and measured pilots
  • Practical experience with AI-assisted tools and agentic workflows, including identifying relevant use cases and integrating AI capabilities into day-to-day engineering practices
  • Strong PostgreSQL skills, including typed relational modelling, JSONB, schema-as-code with migration tooling, and derived views managed as tested transformations
  • Provenance and citation discipline: every extracted fact must be traceable to its source document and passage, and answers must explicitly state when information could not be confirmed
  • Fluent English for written and spoken communication with client teams
  • Будет плюсом: SharePoint and Microsoft Graph APIs or equivalent enterprise content platform integration, bitemporal modelling and document lineage or supersession modelling, graph engines such as Apache AGE or Neo4j, legal, contract, or fund-documentation domain knowledge or ability to learn a document domain in depth, permission-aware retrieval, token-efficiency engineering, exposing capabilities as MCP tools or skills, day-to-day use of AI coding agents, financial services or other regulated on-premise experience, client-facing experience

условия

  • Условий в вакансии нет

Если просят выйти из iCloud, прислать код из SMS, запустить или установить что-то, перевести деньги — не соглашайтесь: это мошенничество.

Про зарплаты

Анонимные данные по зарплатам и грейдам.
Можно сверить вилку с рынком.

Посмотреть зарплаты

Если просят выйти из iCloud, прислать код из SMS, запустить или установить что-то, перевести деньги — не соглашайтесь: это мошенничество.