Чтобы адаптировать резюме под вакансию или составить сопроводительное письмо, загрузи резюме
О рекламодателе
ОБЩЕСТВО С ОГРАНИЧЕННОЙ ОТВЕТСТВЕННОСТЬЮ "ЦЕНТР НАЦИОНАЛЬНЫХ ИНТЕЛЛЕКТУАЛЬНЫХ СИСТЕМ" ИНН: 9704271170
описание
Depending on the role, we might also ask you to do a short presentation, a practical or technical task or have a values focused conversation. We will explain what is involved before anything happens.
CreateFuture is an AI-native consulting partner that works alongside organisations to build digital products and services. Its team develops software, shapes delivery and go-to-market strategies, and creates AI solutions.
задачи
Design and build evaluation pipelines on a platform such as Braintrust, LangSmith, Arize or Weights & Biases, and select the platform that fits the problem
Curate golden datasets representing real user behaviour and design LLM-as-judge and human-in-the-loop scoring pipelines
Build token-level tracing, quality monitoring, hallucination and refusal detection, per-use-case cost attribution, and alerts that catch degradation before users report it
Integrate evaluation gates into CI/CD promotion paths so regressions block releases
Establish baselines, regression suites and drift detection for changes to prompts, models and retrieval
Assess statistical signals, size evaluation sets appropriately, and explain when data cannot answer a question
Present evaluation and observability findings to engineering and product audiences, drive decisions and follow through on resulting changes
Help client teams adopt evaluation-first practices and challenge “ship it and see” approaches with evidence
Plan and prioritise the workstream, estimate accurately, manage changing requirements without compromising quality, and flag timeline risks early
Design evaluation runs and judge models to balance cost and confidence, and justify the approach to budget holders
Work with the client’s existing tools and release process where appropriate, and make the case for changes where needed
Document dataset provenance, scoring rationale and runbooks so client engineers can maintain and extend the evaluation suite
Transfer knowledge to client engineers on evaluation practices, prompt and context engineering, and critical review of model output
Make a clean handover part of the definition of done from the first sprint
требования
Strong, production-grade Python skills and the ability to write maintainable code
Hands-on production experience with Braintrust, LangSmith, Arize, Weights & Biases or an equivalent evaluation and observability platform
Demonstrable experience designing LLM-as-judge and human-in-the-loop pipelines and calibrating them against human judgement
Experience operating LLM-backed features in production, including tracing, latency, token cost, failure modes and retrieval quality
Sufficient fluency with GitHub Actions, GitLab CI or similar to own an evaluation gate in a promotion pipeline
Experience building data pipelines and stores for datasets, traces and results at volume
Ability to communicate technical results to non-technical audiences and tailor the framing to the audience
A track record of becoming productive quickly on unfamiliar programmes and building credibility with engineers
Будет плюсом: experience in regulated industries where model behaviour has compliance or consumer-protection consequences, including iGaming, financial services or health; experience in safer gambling or responsible-messaging contexts; an evaluation framework, write-up or open-source contribution; AWS Certified Machine Learning Engineer (Associate), AWS Certified AI Practitioner, or an associate-level AWS or GCP certification
условия
Contract engagement delivering a defined scope of work on a client programme
Duration, rate and IR35 status will be confirmed at first contact
Travel to client sites or CreateFuture offices may be required as the programme requires
35 Days of leave, including bank holidays
Private medical insurance
Enhanced parental and adoption leave
Financial coaching and a 5% pension match
40 Hours of paid learning and development
Flexible working, with hybrid and remote options
Flexible time management balancing collaboration, client time and focused work