Чтобы адаптировать резюме под вакансию или составить сопроводительное письмо, загрузите резюме
описание
Joi AI is a platform for AI-lationships: personalized, emotionally intelligent connections between people and AI characters. It serves emotional and intimate needs and has tens of millions of conversations. Joi Lab is its open research arm, building open-source generative models, agent architectures, and training infrastructure.
задачи
Speed up and scale production LLM inference using SGLang, KV and prefix caching, batching, quantization, and speculative decoding
Run distributed inference for very large models across multi-GPU and multi-node setups
Benchmark new GPU servers and hardware, bring them into production, and adapt serving code
Lead NLP and CV teams technically by reviewing experiments, setting direction, and addressing wrong turns early
Train and fine-tune language models for AI companions and improve agent harnesses and chat algorithms
Track inference and post-training research and open-source work, and turn findings into the ML roadmap
Collaborate with validation, content, and dataset preparation teams to design experiments and measure model quality
требования
Deep hands-on experience optimizing production LLM inference with SGLang, vLLM, or TensorRT-LLM
Experience with distributed inference or training of large models, including MoE, tensor/expert/pipeline parallelism, and multi-node GPU clusters
Strong understanding of KV cache, attention kernels, batching, quantization, and GPU profiling
Experience training and fine-tuning LLMs, including post-training such as RLHF or DPO
Proven technical leadership through reviews, mentoring, and technical decisions while continuing to write code
Proficiency with PyTorch, transformers, and related libraries
Advanced English or Russian
Будет плюсом: experience at AI-focused startups or companies such as Character AI or OpenAI, backend engineering with Python, Go, or C#, scalable deployment systems, CUDA or Triton kernel development, a computer vision background or experience accelerating generative image or video models, multimodal LLMs, first-author papers or notable open-source contributions, a degree in CS, math, or physics (MSc or PhD)
условия
28 Days of annual leave
7 Yearly wellness days; no sick notes needed
Referral rewards up to $5000
50% Of training, conference, and global meetup costs covered
Subsidized English classes through a corporate discount
Up to $1,000 annually for medical fees or private insurance if not on the group plan
Office equipment or $1000 every three years for a home office or co-working setup
Gratitude bonuses redeemable for swag, massage vouchers, or team adventures