Чтобы адаптировать резюме под вакансию или составить сопроводительное письмо, загрузи резюме
описание
The team is building a sovereign, multi-tenant agentic AI platform and applied products on top of it. Its systems operate under strict data residency constraints and support Arabic and English.
задачи
Set the evaluation strategy, including what to measure, at which layer, with which methodology, and how results inform product decisions
Build shared evaluation infrastructure, including evaluation harnesses, golden-set management, dataset versioning, automated grading, regression detection, and reporting for leadership
Make grading trustworthy through judge model selection, rubric design, calibration against human labels, and identifying when automated grading cannot be trusted
Evaluate retrieval and agents for grounding and citation correctness, tool-use validation, multi-step reasoning, and failure recovery
Develop Arabic golden sets and judges calibrated for Arabic
Run online evaluation and drift detection
Conduct prompt injection, jailbreak, and data-leakage red-teaming
Integrate evaluation and test automation as quality gates in CI/CD
требования
5+ Years of experience building evaluation or test infrastructure used by others, including harnesses, shared libraries, and frameworks
At least 1 year of relevant leadership experience
Expertise in AI evaluation covering output quality, retrieval and grounding, and regression detection for non-deterministic behavior
Proficiency in LLM-as-judge methodology, including rubric design, calibration against human labels, and understanding its failure modes
Advanced Python proficiency sufficient to build shared libraries
Intermediate competency in TypeScript or Java
Knowledge of statistics for non-deterministic systems, including sampling, confidence intervals, inter-rater agreement, and significance
Skills in CI/CD framework design across API, web, and data surfaces, including test selection, parallelization, and flake management
Background in golden sets and dataset versioning
English proficiency at Upper-Intermediate level (B2) or higher
Будет плюсом: familiarity with ragas, DeepEval, Promptfoo, Braintrust, LangSmith, or Langfuse; expertise in Arabic evaluation, including golden sets, dialect coverage, right-to-left validation, and judge calibration; background in adversarial and security testing, including red-teaming; capability to perform voice and conversational evaluation and performance testing against AI services; experience delivering in a regulated or government environment