Если вы раньше входили через Google, сбросьте пароль для своей Gmail-почты через кнопку «Забыли пароль?» на экране входа. Затем войдите по email и новому паролю.
Если аккаунта ещё нет, зарегистрируйтесь с Gmail-почтой — после подтверждения почты мы предложим задать пароль.
Что нового
Загружаю обновления...
Что нового
Загружаю обновления...
Работа найдется быстрее с подпискойКандидат найдётся быстрее с подпиской
Чтобы адаптировать резюме под вакансию или составить сопроводительное письмо, загрузите резюме
описание
Must be available to work from 13:00 to 21:00 Polish time.
EPAM provides digital engineering, cloud, and AI-enabled transformation services, as well as digital product design and consulting.
задачи
Research and evolve automation frameworks in line with Gen AI tooling and best practices;
Design and automate evaluation of Gen AI features including grounding, answer accuracy, determinism, reproducibility, precision, recall, and criteria recall;
Build automated LLM test harnesses that scale evaluation beyond human-in-the-loop;
Select and apply Gen AI evaluation frameworks to measure answer quality and pipeline efficiency;
Perform manual testing to validate new features, integrations, and user stories;
Build and maintain test cases from requirements and user stories;
Test applications including AI agents, APIs, databases, and other integrations;
Collaborate with product, engineering, and operations teams to understand requirements and deployment environments;
Track and report test results, defects, and quality metrics;
Assist with troubleshooting production issues and escalate risks;
Guide and support team members, including onshore and offshore consultants.
требования
5+ Years of experience in software QA, with at least 1 year focused on testing AI agents, agentic solutions or LLM-based systems;
Hands-on experience with both manual and automated testing of AI agents, including prompt/instruction testing and evaluation of agentic workflows;
Strong programming skills in Python test automation using pytest or equivalent, scripting, and AI/ML library integration;
Expertise in AI agent frameworks, prompt engineering, and evaluation metrics for LLM-based systems;
Demonstrated experience testing and evaluating Gen AI / LLM applications including grounding, answer accuracy, and hallucination/determinism checks;
Applied knowledge of Gen AI / LLM evaluation frameworks and metrics such as precision, recall, criteria recall, and efficiency;
Familiarity with issue and test management tools such as Jira, QMetry, and TestRail;
Experience with version control systems and integrating tests into CI/CD pipelines;
Flexibility to use AI-powered tools for QA such as GitHub Copilot and LLM-based test generation;
Understanding of cloud environments, particularly AWS;
Excellent communication, collaboration, and leadership skills;
Nice to have: Experience with agentic AI platforms such as LangChain, OpenAI Function Calling or similar, experience with AI safety, bias and reliability testing, experience with test data generation for AI/ML systems.