вчера

machine learning engineer for speech systems

выше рынка на 466,0%
вакансия 1 443 095 ₽
в среднем 254 966 ₽
Загрузи резюме, чтобы видеть мэтчи с вакансией

подготовьтесь к отклику

ai-инструменты

Чтобы адаптировать резюме под вакансию или составить сопроводительное письмо, загрузите резюме

описание

Cantina is a social platform with an AI character creator. Its lifelike AI characters interact across voice, video, and text, enabling personalized content and group chat.

задачи

  • Architect, implement, pre-train, fine-tune, and post-train large-scale speech models;
  • Independently lead small research projects and collaborate on larger team initiatives;
  • Design, run, and analyze scientific experiments;
  • Develop and improve development tooling;
  • Contribute across the stack, from low-level optimizations to high-level model design;
  • Define data requirements and collaborate on acquisition, curation, augmentation, labeling quality, and synthetic data strategies;
  • Design automated objective and subjective evaluations, including listening tests, SV/WER/ASR-based metrics, robustness and bias checks, and red-team studies;
  • Harden the training, evaluation, and inference pipeline;
  • Profile latency, memory, and cost, and meet production SLAs with monitoring and rollback;
  • Partner with infrastructure on distributed training and inference on cloud fleets;
  • Productionize models with reliability and observability;
  • Contribute to safety and consent guardrails and misuse and abuse mitigation.

требования

  • Exceptional research and development experience with large-scale audio models, including models over 3B parameters and over 500k hours of data;
  • Exceptional understanding and hands-on experience with transformer architectures, diffusion models including distillation and streaming, or audio language modeling;
  • Strong experience with multi-node and multi-GPU distributed model training;
  • Strong software engineering skills and experience building complex systems;
  • Strong PyTorch and performance engineering skills, including profiling and CUDA/Triton/C++ as needed;
  • Experience writing reliable, production-quality code;
  • Experience shipping large-scale speech and audio models to production;
  • Background working with large-scale ML data;
  • Ability to iterate on data and assess quality using subjective and objective signals;
  • Notable publications or open-source contributions in speech, audio, or ML;
  • Experience with voice cloning, speech control, or voice generation;
  • Nice to have: Experience shipping large-scale TTS, VC, or ASR models to production, work on large-scale ML systems, experience with audio language modeling and transformer architectures, experience with voice cloning, speech control, or voice generation, background in processing large-scale ML data, publications or notable open-source contributions in speech, audio, or ML.

условия

  • Annual base salary range: $200,000-$220,000 (€170,000-€190,000);
  • Competitive salary and generous company equity;
  • Medical, dental, and vision insurance with 99.99% of premiums covered by Cantina;
  • 42 Days of paid time off, including 15 PTO days, 10 sick days, 15 company holidays, and 2 floating holidays;
  • Generous parental leave and fertility support;
  • 401(K) retirement savings plan;
  • Lifestyle spending account of $500/month;
  • Complimentary lunch and snacks for in-office employees;
  • One Medical membership.

Если просят выйти из iCloud, прислать код из SMS, запустить или установить что-то, перевести деньги — не соглашайтесь: это мошенничество.

Про зарплаты

Анонимные данные по зарплатам и грейдам.
Можно сверить вилку с рынком.

Посмотреть зарплаты

Если просят выйти из iCloud, прислать код из SMS, запустить или установить что-то, перевести деньги — не соглашайтесь: это мошенничество.