4 окт

platform engineer for AI evaluation

ориентир по рынку
вакансия зп не указана
в среднем 353 643 ₽
Загрузи резюме, чтобы видеть мэтчи с вакансией

подготовься к отклику

ai-инструменты

Чтобы адаптировать резюме под вакансию или составить сопроводительное письмо, загрузи резюме

Рекламный баннер: ОБЩЕСТВО С ОГРАНИЧЕННОЙ ОТВЕТСТВЕННОСТЬЮ "ЦЕНТР НАЦИОНАЛЬНЫХ ИНТЕЛЛЕКТУАЛЬНЫХ СИСТЕМ"
О рекламодателе
ОБЩЕСТВО С ОГРАНИЧЕННОЙ ОТВЕТСТВЕННОСТЬЮ "ЦЕНТР НАЦИОНАЛЬНЫХ ИНТЕЛЛЕКТУАЛЬНЫХ СИСТЕМ"
ИНН: 9704271170

описание

Depending on the role, we might also ask you to do a short presentation, a practical or technical task or have a values focused conversation. We will explain what is involved before anything happens.

CreateFuture is an AI-native consulting partner that works alongside organisations to build digital products and services. Its team develops software, shapes delivery and go-to-market strategies, and creates AI solutions.

задачи

  • Design and build evaluation pipelines on a platform such as Braintrust, LangSmith, Arize or Weights & Biases, and select the platform that fits the problem
  • Curate golden datasets representing real user behaviour and design LLM-as-judge and human-in-the-loop scoring pipelines
  • Build token-level tracing, quality monitoring, hallucination and refusal detection, per-use-case cost attribution, and alerts that catch degradation before users report it
  • Integrate evaluation gates into CI/CD promotion paths so regressions block releases
  • Establish baselines, regression suites and drift detection for changes to prompts, models and retrieval
  • Assess statistical signals, size evaluation sets appropriately, and explain when data cannot answer a question
  • Present evaluation and observability findings to engineering and product audiences, drive decisions and follow through on resulting changes
  • Help client teams adopt evaluation-first practices and challenge “ship it and see” approaches with evidence
  • Plan and prioritise the workstream, estimate accurately, manage changing requirements without compromising quality, and flag timeline risks early
  • Design evaluation runs and judge models to balance cost and confidence, and justify the approach to budget holders
  • Work with the client’s existing tools and release process where appropriate, and make the case for changes where needed
  • Document dataset provenance, scoring rationale and runbooks so client engineers can maintain and extend the evaluation suite
  • Transfer knowledge to client engineers on evaluation practices, prompt and context engineering, and critical review of model output
  • Make a clean handover part of the definition of done from the first sprint

требования

  • Strong, production-grade Python skills and the ability to write maintainable code
  • Hands-on production experience with Braintrust, LangSmith, Arize, Weights & Biases or an equivalent evaluation and observability platform
  • Demonstrable experience designing LLM-as-judge and human-in-the-loop pipelines and calibrating them against human judgement
  • Experience operating LLM-backed features in production, including tracing, latency, token cost, failure modes and retrieval quality
  • Sufficient fluency with GitHub Actions, GitLab CI or similar to own an evaluation gate in a promotion pipeline
  • Experience building data pipelines and stores for datasets, traces and results at volume
  • Ability to communicate technical results to non-technical audiences and tailor the framing to the audience
  • A track record of becoming productive quickly on unfamiliar programmes and building credibility with engineers
  • Будет плюсом: experience in regulated industries where model behaviour has compliance or consumer-protection consequences, including iGaming, financial services or health; experience in safer gambling or responsible-messaging contexts; an evaluation framework, write-up or open-source contribution; AWS Certified Machine Learning Engineer (Associate), AWS Certified AI Practitioner, or an associate-level AWS or GCP certification

условия

  • Contract engagement delivering a defined scope of work on a client programme
  • Duration, rate and IR35 status will be confirmed at first contact
  • Travel to client sites or CreateFuture offices may be required as the programme requires
  • 35 Days of leave, including bank holidays
  • Private medical insurance
  • Enhanced parental and adoption leave
  • Financial coaching and a 5% pension match
  • 40 Hours of paid learning and development
  • Flexible working, with hybrid and remote options
  • Flexible time management balancing collaboration, client time and focused work

Если просят выйти из iCloud, прислать код из SMS, запустить или установить что-то, перевести деньги — не соглашайся: это мошенничество.

Про зарплаты

Анонимные данные по зарплатам и грейдам.
Можно сверить вилку с рынком.

Посмотреть зарплаты

Если просят выйти из iCloud, прислать код из SMS, запустить или установить что-то, перевести деньги — не соглашайся: это мошенничество.