Чтобы адаптировать резюме под вакансию или составить сопроводительное письмо, загрузи резюме
О рекламодателе
ОБЩЕСТВО С ОГРАНИЧЕННОЙ ОТВЕТСТВЕННОСТЬЮ "ЦЕНТР НАЦИОНАЛЬНЫХ ИНТЕЛЛЕКТУАЛЬНЫХ СИСТЕМ" ИНН: 9704271170
описание
Toloka AI creates data that powers leading GenAI models and innovations. It combines experts, crowdsourcing, and a technology platform to teach AI models to reason and evaluate their efficacy and safety.
задачи
Define the multi-quarter strategy and architecture for Toloka’s post-training stack, keeping fine-tuning, RL, and evaluation capabilities ahead of industry trends
Architect advanced post-training and RL pipelines, including GRPO, PPO/DPO, reward modeling, Process Reward Models, and RLAIF, for platform-wide LLM alignment and agent behavior
Design and calibrate automated evaluation ecosystems, including human-aligned LLM-as-a-judge frameworks, dynamic benchmarking suites, and regression control setups
Lead research on model distillation, quantization, context compression, and speculative decoding to optimize latency-cost trade-offs at platform scale
Architect autonomous guiding agents and tool-use workflows, establishing frameworks for multi-step reasoning, self-correction, and evaluation-driven feedback loops
Raise the team’s engineering and research standards through architectural design reviews, hands-on mentoring of Senior ML Researchers, and reproducible ML R&D best practices
Partner with Product and Engineering directors to translate client challenges into scalable platform architecture
Author research write-ups and blog posts, and deliver external tech talks
требования
6+ Years in ML engineering or applied research, including 3+ years leading LLM post-training, fine-tuning, or alignment initiatives at scale
Deep theoretical and practical mastery of LLM alignment: SFT, LoRA/PEFT, RLHF/RLAIF (GRPO, DPO, PPO), reward modeling, and reasoning-oriented post-training
Proven experience building evaluation pipelines from scratch, addressing LLM-as-a-judge biases, and calibrating automated evaluations against human gold standards
Strong proficiency in Python, PyTorch, distributed training frameworks (Deepspeed, Megatron, FSDP), and high-performance inference engines (vLLM, TensorRT-LLM)
Ability to turn ambiguous technical challenges into robust platform features while balancing advanced research, production reliability, and cost constraints
Proven experience as a technical leader, mentoring senior researchers or engineers and influencing cross-functional roadmaps
Fluent spoken and written English (C1), with the ability to explain complex technical concepts to internal teams, external clients, and the research community
условия
Competitive compensation package including base salary, bonus, and ESOP