Lead Research Engineer / Research Scientist - Post-Training, Reinforcement Learning & Training Systems in Lausanne

Lead Research Engineer / Research Scientist - Post-Training, Reinforcement Learning & Training Systems in Lausanne

Lausanne Vollzeit Vor Ort
G

ppGiotto.ai is a Switzerland-based AI company building intelligence systems for Switzerland and Europe.br/Our mission is to enable governments and enterprises to retain control over the AI systems they use without compromising access to advanced reasoning capabilities. Giotto combines portable, configurable models with an AI operating system, integrating open and proprietary weights, datasets, tools, and deployment components.br/ /p h3About the role /h3 pWe are looking for a Lead Research Engineer or Research Scientist to own and lead the training and optimisation side of our complete post-training stack. Starting from pretrained checkpoints, you will design, implement, scale, and operate the methods required to produce capable, reliable, and controllable production models. Your scope will include supervised fine-tuning, preference optimisation, reinforcement learning, reward and verifier integration, policy distillation or consolidation, and distributed training. You will be expected to set technical direction in these areas, make key design decisions, and help guide the work of other researchers and engineers while remaining deeply hands‑on. This is not a single‑GPU fine‑tuning or adapter‑only role. You should be comfortable operating training workloads where memory, communication, rollout generation, hardware topology, and fault recovery must be designed together. /p h3You will: /h3 ul liOwn the end-to-end post‑training pipeline from pretrained checkpoint to production candidate. /li liSet technical direction and priorities for post‑training and reinforcement‑learning work. /li liDesign and execute full‑parameter and parameter‑efficient SFT. /li liImplement preference optimisation, RLHF, RLAIF, reinforcement learning with verifiable rewards, and related methods. /li liDevelop training strategies for reasoning, coding, tool use, multilingual behaviour, and long‑horizon agent tasks. /li liIntegrate reward models, verifiers, critics, graders, and process‑ or outcome‑based rewards. /li liBuild scalable rollout‑generation systems for iterative and on‑policy training. /li liDesign multi‑stage curricula combining SFT, reinforcement learning, rejection sampling, distillation, and policy consolidation. /li liScale training across multiple machines and accelerators using appropriate combinations of data, tensor, pipeline, sequence, context, or expert parallelism. /li liSelect sharding, precision, checkpointing, optimiser, batch‑size, sequence‑length, and activation‑recomputation strategies. /li liEstimate memory, communication, throughput, rollout capacity, and compute requirements before launching major runs. /li liProfile and improve accelerator utilisation, communication efficiency, data loading, and end‑to‑end training time. /li liDiagnose numerical instability, communication failures, out‑of‑memory errors, stragglers, checkpoint issues, and convergence regressions. /li liInvestigate reward hacking, entropy collapse, KL drift, stale rollouts, mode collapse, grader exploitation, and benchmark overfitting. /li liBuild reliable checkpointing, recovery, monitoring, and reproducibility procedures. /li liCollaborate closely with data, evaluation, infrastructure, and inference teams. /li liProvide technical guidance and mentorship to other researchers and engineers working on training and post‑training. /li liContribute clean, tested code, technical reports, and operational runbooks. /li /ul h3We are looking for demonstrated experience in most of the following areas: /h3 ul liOwnership of large‑scale language‑model training or post‑training runs across multiple machines and accelerators. /li liExperience with workloads for which straightforward single‑node training or pure data parallelism was insufficient. /li liDeep proficiency with Python, PyTorch, autograd, mixed precision, optimisation, and distributed execution. /li liPractical experience with PyTorch Distributed, FSDP, DeepSpeed, Megatron‑Core, or an equivalent framework. /li liAbility to select parallelism and sharding strategies based on model, sequence, memory, and network constraints. /li liStrong understanding of SFT, preference optimisation, reinforcement learning, reward modelling, KL regularisation, sampling, and training stability. /li liExperience operating high‑throughput inference or rollout systems as part of a training loop. /li liAbility to debug across model code, distributed communication, numerical optimisation, data, and infrastructure. /li liStrong experimental design and the ability to distinguish algorithmic improvements from evaluation or systems artefacts. /li liExperience building reliable, observable, and reproducible research software. /li liPersonal ownership of consequential decisions affecting a substantial training programme. /li liDemonstrated ability to provide technical direction, review complex design choices, and raise the technical level of a team. /li /ul pA PhD is not required. We value exceptional technical work, strong judgement, and demonstrated ownership.br/ /p h3Relevant stack: /h3 ul liPython and PyTorch. /li liPyTorch Distributed and FSDP. /li liDeepSpeed, Megatron‑Core, or comparable frameworks. /li liHugging Face Transformers. /li liCUDA and NCCL. /li livLLM, SGLang, or similar rollout engines. /li liRay, Slurm, Kubernetes, or comparable orchestration systems. /li liMLflow or Weights Biases. /li liDocker, GCP, GitLab CI, profiling, monitoring, and pytest. /li /ul pExperience with CUDA or Triton, long-context training, sparse models, asynchronous RL, stateful agent environments, distillation, or deployment‑aware post‑training would be especially valuable. /p h3You may be a strong fit if you: /h3 ul liEnjoy working at the intersection of model research and distributed systems. /li liCan move from paper reproduction to reliable scaled implementation. /li liAre comfortable taking responsibility for expensive and operationally demanding experiments. /li liApproach failures methodically across algorithms, data, numerical stability, and infrastructure. /li liCare about held‑out capability and reliability, not only training loss or reward. /li liAre comfortable setting technical direction while remaining deeply hands‑on. /li liWant meaningful ownership of a complete model programme. /li /ul h3Location and work style: /h3 pWe offer full‑time employment in Switzerland. /p ul liRemote work is supported. /li liThe team gathers approximately twice per month in a Swiss office. /li liExceptional candidates elsewhere in Europe may be considered. /li /ul /p #J-18808-Ljbffr

Lead Research Engineer / Research Scientist - Post-Training, Reinforcement Learning & Training Systems in Lausanne Arbeitgeber: Giotto.ai SA

Giotto.ai ist ein hervorragender Arbeitgeber, der seinen Mitarbeitern die Möglichkeit bietet, an der Spitze der KI-Technologie zu arbeiten und dabei die strategische Unabhängigkeit Europas zu fördern. Mit einem hybriden Arbeitsmodell, das Remote-Arbeit unterstützt und regelmäßige Teamtreffen in der Schweiz ermöglicht, fördert das Unternehmen eine dynamische und kollaborative Arbeitskultur. Zudem legt Giotto.ai großen Wert auf die persönliche und berufliche Weiterentwicklung seiner Mitarbeiter, indem es ihnen Zugang zu innovativen Projekten und wertvollen Schulungen bietet.

G

Kontaktdaten:

Giotto.ai SA Recruiting-Team