platform engineer for AI infrastructure
генерация резюме под вакансию
сопроводительное письмо
описание
Langdock is an AI platform used by more than 10,000 companies to provide secure access to leading AI models, build and share agents, and automate repetitive workflows. The platform is designed for European enterprises while preserving control over data, model providers, and deployment environments.
задачи
- Own shared platform capabilities as internal products, including their interfaces, reliability, evolution, and long-term operation;
- Turn model access, agent execution, authorization, queues, integrations, document processing, and sandboxed compute into stable, reusable building blocks;
- Own services from design through production operation;
- Define service APIs and failure behavior;
- Plan migrations, implement and deploy services, make them observable, and respond to production failures;
- Work with Systems Engineers on underlying mechanisms while remaining accountable for service boundaries and lifecycles;
- Develop a durable agent runtime that supports approvals, external events, process recovery, deterministic context reconstruction, compaction, and prompt-cache efficiency;
- Build the Model Gateway with provider adapters, routing, rate limits, bounded fallback, streaming safety, usage tracking, and a control plane;
- Move selected capabilities from the TypeScript monolith into independently operated services with Protobuf and gRPC contracts;
- Develop consistent authentication, token refresh, rate limits, approvals, outbound network safety, and tool execution across REST APIs, MCP, and A2A;
- Define the execution API, lifecycle, permission model, resource budgets, and operational boundary for shared code execution;
- Build load testing for HTTP, gRPC, and streaming workloads;
- Implement structured logging without user content or PII;
- Perform online database changes for large PostgreSQL tables;
- Own a platform area’s interface and production operation;
- Own shipped changes in production and lead fixes when incidents occur;
- Contribute to roadmap, prioritization, technical decisions, and production operations within a squad;
- Use design documents to define architectural boundaries, tradeoffs, failure modes, migrations, and rollouts;
- Invest in observability, migration tooling, automated recovery, and runbooks.
требования
- Several years of experience, typically 3 to 6, building and operating backend systems under real load;
- Strong TypeScript and Node.js skills;
- Comfortable with Kubernetes, Terraform, networking, databases, and queues;
- Strong platform design judgment;
- Ability to define stable service boundaries and APIs;
- Understanding of data integrity, tenant isolation, backward compatibility, partial failure, and safe rollout;
- Ability to investigate unfamiliar components and failure modes;
- Ability to make systems observable and respond to production failures;
- Ability to turn incidents into lasting improvements;
- Responsible use of AI tools while owning architecture and shipped changes;
- Ability to share context, ask for input, disagree directly and respectfully, and help other engineers succeed;
- Nice to have: Go experience.
условия
- All roles include equity;
- Requires an existing right to work in Germany; visa sponsorship is not currently provided;
- Engineers usually start around 8:30;
- Lunch is usually provided, and dinner is available for people who stay later;
- Running and gym activities are part of the routine for many employees.
навыки
Если просят войти через iCloud, отправить коды из SMS, запустить код, что-то установить, перевести деньги или сделать что угодно, связанное с деньгами, не соглашайтесь: это признаки мошенничества.