7 окт

ML researcher

ориентир по рынку
вакансия зп не указана
в среднем 305 081 ₽
Загрузи резюме, чтобы видеть мэтчи с вакансией

подготовься к отклику

ai-инструменты

Чтобы адаптировать резюме под вакансию или составить сопроводительное письмо, загрузи резюме

Рекламный баннер: ОБЩЕСТВО С ОГРАНИЧЕННОЙ ОТВЕТСТВЕННОСТЬЮ "ЦЕНТР НАЦИОНАЛЬНЫХ ИНТЕЛЛЕКТУАЛЬНЫХ СИСТЕМ"
О рекламодателе
ОБЩЕСТВО С ОГРАНИЧЕННОЙ ОТВЕТСТВЕННОСТЬЮ "ЦЕНТР НАЦИОНАЛЬНЫХ ИНТЕЛЛЕКТУАЛЬНЫХ СИСТЕМ"
ИНН: 9704271170

описание

Toloka AI creates data that powers leading GenAI models and innovations. It combines experts, crowdsourcing, and a technology platform to teach AI models to reason and evaluate their efficacy and safety.

задачи

  • Define the multi-quarter strategy and architecture for Toloka’s post-training stack, keeping fine-tuning, RL, and evaluation capabilities ahead of industry trends
  • Architect advanced post-training and RL pipelines, including GRPO, PPO/DPO, reward modeling, Process Reward Models, and RLAIF, for platform-wide LLM alignment and agent behavior
  • Design and calibrate automated evaluation ecosystems, including human-aligned LLM-as-a-judge frameworks, dynamic benchmarking suites, and regression control setups
  • Lead research on model distillation, quantization, context compression, and speculative decoding to optimize latency-cost trade-offs at platform scale
  • Architect autonomous guiding agents and tool-use workflows, establishing frameworks for multi-step reasoning, self-correction, and evaluation-driven feedback loops
  • Raise the team’s engineering and research standards through architectural design reviews, hands-on mentoring of Senior ML Researchers, and reproducible ML R&D best practices
  • Partner with Product and Engineering directors to translate client challenges into scalable platform architecture
  • Author research write-ups and blog posts, and deliver external tech talks

требования

  • 6+ Years in ML engineering or applied research, including 3+ years leading LLM post-training, fine-tuning, or alignment initiatives at scale
  • Deep theoretical and practical mastery of LLM alignment: SFT, LoRA/PEFT, RLHF/RLAIF (GRPO, DPO, PPO), reward modeling, and reasoning-oriented post-training
  • Proven experience building evaluation pipelines from scratch, addressing LLM-as-a-judge biases, and calibrating automated evaluations against human gold standards
  • Strong proficiency in Python, PyTorch, distributed training frameworks (Deepspeed, Megatron, FSDP), and high-performance inference engines (vLLM, TensorRT-LLM)
  • Ability to turn ambiguous technical challenges into robust platform features while balancing advanced research, production reliability, and cost constraints
  • Proven experience as a technical leader, mentoring senior researchers or engineers and influencing cross-functional roadmaps
  • Fluent spoken and written English (C1), with the ability to explain complex technical concepts to internal teams, external clients, and the research community

условия

  • Competitive compensation package including base salary, bonus, and ESOP
  • Paid PTO and benefits vary by location
  • IT setup and home office allowances

Платформа данных для обучения ИИ-моделей.

Если просят выйти из iCloud, прислать код из SMS, запустить или установить что-то, перевести деньги — не соглашайся: это мошенничество.

Спроси Хайрика про вакансию

Сверит с твоим резюме, подскажет вилку и вопросы на собесе.

Если просят выйти из iCloud, прислать код из SMS, запустить или установить что-то, перевести деньги — не соглашайся: это мошенничество.