Чтобы адаптировать резюме под вакансию или составить сопроводительное письмо, загрузите резюме
описание
The lab builds resettable, interactive digital environments where frontier AI agents learn by actively operating real software. It develops data infrastructure that enables models to train reliably on large, complex real-world datasets and collaborates with leading AI research labs.
задачи
Architect, build, and operate scalable distributed systems to acquire, ingest, and process large, unstructured datasets for agent training environments
Design and implement automated data quality, validation, sanitisation, and PII-handling pipelines resilient to adversarial and messy real-world data
Make foundational trade-offs across storage tiers, query performance, reliability, observability, and infrastructure cost without being tied to a fixed vendor stack
Partner with platform, software, and AI research teams to establish platform-wide architectural patterns
Own core data systems from initial problem definition through production operations
Write production code, debug complex distributed systems across platform and AI engineering boundaries, and design validation, sanitisation, and ingestion infrastructure for agent training
требования
Proven track record of owning substantial production data systems end to end, with personal accountability for architecture, implementation, and operational outcomes rather than just configuring existing ETL tools
Strong, recent hands-on software engineering fundamentals and ability to work in unfamiliar codebases and debug distributed systems failure modes
Demonstrated experience tackling large, complex, or messy datasets and articulating failure modes, trade-offs, and architectural alternatives
Comfortable operating in high ambiguity: taking an unformed problem, deciding what to build, and executing without a detailed specification
Based in or willing to work with engineering teams in London or Paris