Если вы раньше входили через Google, сбросьте пароль для своей Gmail-почты через кнопку «Забыли пароль?» на экране входа. Затем войдите по email и новому паролю.
Если аккаунта ещё нет, зарегистрируйтесь с Gmail-почтой, после подтверждения почты мы предложим задать пароль.
Что нового
Загружаю обновления...
Что нового
Загружаю обновления...
Работа найдется быстрее с подпискойКандидат найдётся быстрее с подпиской
Чтобы адаптировать резюме под вакансию или составить сопроводительное письмо, загрузите резюме
описание
Please Note: Due to the high volume of applications, only shortlisted candidates will be contacted.
How would you rate your hands-on experience with terminal-based software engineering workflows (e.g. compiling code, debugging failing tests, managing environments, navigating repositories, fixing dependency issues)? * • Which of the following tasks have you personally performed while evaluating or reviewing AI-generated code? (Select all that apply) * • What best describes your experience with AI data production, RLHF, LLM evaluation, or annotation projects? * • How would you describe your level of hands-on expertise in your primary programming language (e.g. Python, Go, JavaScript/TypeScript, Java, C++, etc.)? *
Kake is an expert human data platform that supports AI agents and LLM development by providing high-quality training data and human evaluation.
задачи
Create and review coding tasks based on real-world software engineering scenarios, including debugging, refactoring, code generation, API usage, automated tests, performance, security, and edge cases;
Write correct, clear, testable reference solutions aligned with task requirements;
Evaluate AI-generated code and responses using structured rubrics for correctness, clarity, security, performance, maintainability, and instruction-following;
Compare multiple model responses, select the strongest answer, and justify the decision with clear technical reasoning;
Identify bugs, hallucinated APIs, missing edge cases, weak explanations, and poor engineering decisions in AI-generated outputs;
Work with terminal-based development workflows, including running tests, debugging issues, managing dependencies, and navigating repositories;
Follow detailed guidelines consistently and participate in calibration activities to ensure high-quality, reliable evaluations.
требования
5+ Years of professional software engineering experience in a backend, fullstack, or systems role;
Strong proficiency in at least one core programming language, ideally Python, JavaScript/TypeScript, Go, Java, C++, or SQL;
Hands-on experience with Terminal-Bench and evaluating AI agent performance on terminal-based tasks;
Comfortable working with Git, command line/terminal, and common development workflows;
Ability to evaluate code critically for design quality, security, and maintainability;
Prior experience in AI data production, RLHF, data annotation, or LLM evaluation projects;
Excellent written and verbal communication skills in English;
Ability to work independently in a remote, asynchronous, fast-paced environment;
High attention to detail and ability to follow complex, rubric-based guidelines consistently;
Nice to have: Experience with Python-heavy workflows, automated testing frameworks, Docker, Linux, bash, containerized environments, repo-level code reasoning, large codebases, open-source contributions, backend systems, data engineering, DevOps, infrastructure, or security.