вчера

ai engineer for high-performance inference

ориентир по рынку
вакансия зп не указана
в среднем 389 693 ₽
Загрузи резюме, чтобы видеть мэтчи с вакансией

подготовьтесь к отклику

ai-инструменты

Чтобы адаптировать резюме под вакансию или составить сопроводительное письмо, загрузите резюме

описание

FAR Labs is a unified platform for high-performance AI inference. It provides dedicated Inference and Model APIs with flexible workload placement across its infrastructure, a customer’s cloud or hardware, and shared GPUs.

задачи

  • Set the technical direction for the inference stack and benchmarking practices;
  • Improve the speed and efficiency of model serving;
  • Validate performance gains with credible benchmarks;
  • Lead the engineers building the platform;
  • Shape the technology, performance standards, and engineering culture behind FAR Labs.

требования

  • Production experience with LLM inference-serving systems;
  • Expertise with vLLM, SGLang, TensorRT-LLM, or similar stacks;
  • Strong knowledge of batching, paged attention, KV-cache management, and modern serving architectures;
  • GPU performance engineering experience with CUDA or Triton;
  • Solid understanding of GPU utilisation, memory bandwidth, and precisions such as FP8 and FP4;
  • English B2.

условия

  • No conditions specified

Если просят выйти из iCloud, прислать код из SMS, запустить или установить что-то, перевести деньги — не соглашайтесь: это мошенничество.

Про зарплаты

Анонимные данные по зарплатам и грейдам.
Можно сверить вилку с рынком.

Посмотреть зарплаты

Если просят выйти из iCloud, прислать код из SMS, запустить или установить что-то, перевести деньги — не соглашайтесь: это мошенничество.