13 авг

qa engineer (auto) for AI evaluation systems

ориентир по рынку
вакансия зп не указана
в среднем 227 573 ₽
Загрузи резюме, чтобы видеть мэтчи с вакансией

подготовьтесь к отклику

ai-инструменты

Чтобы адаптировать резюме под вакансию или составить сопроводительное письмо, загрузите резюме

описание

EPAM provides enterprise software products, open source solutions, and accelerators, including AI-focused solutions and services. The AI Platform team develops the Agent Evaluation Framework built on AWS AgentCore for enterprise-grade agentic AI infrastructure.

задачи

  • Design and implement a functional test suite for the AWS AgentCore Evaluation API using pytest, covering known-good and known-bad session pair validation, multi-evaluator execution correctness, and edge cases;
  • Develop and maintain integration tests for on-demand evaluation mode and integrate them into the CI/CD pipeline with automated execution on each build;
  • Validate online-mode sampling accuracy, design test scenarios, define acceptance criteria, and report deviations with reproducible evidence;
  • Conduct and document a feasibility assessment for non-AgentCore runtime evaluation by analyzing alternative runtimes, defining evaluation methodology, and delivering a structured findings report;
  • Test OpenTelemetry trace-based evaluation inputs and validate ADOT trace ingestion, trace structure correctness, and evaluator input integrity;
  • Collaborate with platform engineers to clarify evaluation contracts, reproduce defects, and align on quality gates;
  • Maintain test documentation, including test plans, test reports, defect logs, and evaluation feasibility artifacts in Confluence and Jira;
  • Participate in Agile ceremonies, including sprint planning, daily standups, demos, and retrospectives;
  • Contribute to EngX practices through code review of test scripts, CI/CD pipeline integration, and test coverage reporting.

требования

  • 5+ Years of production experience in QA automation or ML/AI system testing;
  • Proven experience testing AI/LLM systems, including evaluation pipelines, model outputs, or agent behavior validation;
  • Proficiency in Python test automation, including pytest, fixtures, parametrize, and mocking with unittest.mock and moto;
  • Knowledge of the AWS AgentCore Evaluation API, including on-demand and online evaluation modes;
  • Familiarity with OpenTelemetry and ADOT for trace-based evaluation input testing and trace structure validation;
  • Skills in REST API testing, including request/response validation and authentication with SigV4 and bearer tokens;
  • Experience integrating tests with CI/CD using GitHub Actions, Jenkins, or equivalent tools, including test pipeline configuration;
  • Background in test data management, including known-good and known-bad session pair design and synthetic trace generation;
  • Expertise in functional and integration test design for AI/ML evaluation pipelines;
  • Competency in defect lifecycle management, including Jira, reproducible bug reports, and root cause analysis;
  • Understanding of LLM and agent evaluation concepts, including correctness scoring, sampling strategies, and evaluator chaining;
  • Ability to work independently after onboarding, manage own tasks, report status, and proactively escalate blockers;
  • Strong analytical skills for defining test scenarios from ambiguous or evolving specifications;
  • English B2+ level, written and verbal, for daily collaboration with distributed teams;
  • Nice to have: hands-on experience with AWS AgentCore Evaluation API or AWS Bedrock testing, experience testing OpenTelemetry or distributed tracing pipelines, familiarity with multi-evaluator execution patterns and correctness validation strategies, experience writing feasibility assessments or technical reports for stakeholders, knowledge of agentic AI frameworks such as LangGraph and Strands Agents, ISTQB CT-AI certification or equivalent AI testing qualification, experience with AI Ready or AI Practitioner practices at EPAM.

условия

  • Remote work from Poland.

Если просят выйти из iCloud, прислать код из SMS, запустить или установить что-то, перевести деньги — не соглашайтесь: это мошенничество.

Про зарплаты

Анонимные данные по зарплатам и грейдам.
Можно сверить вилку с рынком.

Посмотреть зарплаты

Если просят выйти из iCloud, прислать код из SMS, запустить или установить что-то, перевести деньги — не соглашайтесь: это мошенничество.