Чтобы адаптировать резюме под вакансию или составить сопроводительное письмо, загрузи резюме
описание
The company is building an AI agent that performs tasks for everyday users, runs errands, manages workflows, and maintains context across long, complex conversations.
задачи
Build end-to-end pipelines for data, training, evaluation, and inference
Adapt and fine-tune models using LoRA, QLoRA, SFT, DPO, and distillation
Architect inference systems that meet latency and cost constraints
Create pipelines for high-quality synthetic and real-world training data
Evaluate robustness, safety, bias, and production behavior beyond benchmarks
Own deployment, including GPU optimization, quantization, memory efficiency, and scaling
Work with application engineers to integrate ML into backend, mobile, and desktop applications
Take ownership of ambiguous problems from zero to one
Ship, iterate, and learn from production
требования
Deep understanding of deep learning and transformer architectures
Proven experience training, fine-tuning, or shipping large-scale models in production
Strong with at least one major ML framework, such as PyTorch or JAX, and able to pick up others quickly
Familiarity with distributed training and inference tools, including DeepSpeed, FSDP, Megatron, ZeRO, or Ray
Engineering discipline and ability to write readable, robust, maintainable code
Experience optimizing for GPU constraints, including quantization, mixed precision, and memory
Будет плюсом: LLM inference frameworks such as vLLM, TensorRT-LLM, or FasterTransformer; RLHF methods such as PPO, DPO, or ORPO; open-source contributions to ML or systems libraries; scientific computing, compiler, or GPU kernel experience; multimodal or diffusion model background; large-scale data processing with Arrow, Spark, or Ray