Overview
As Senior Research Scientist at DeepL, you will lead the direction of steerable, high‑quality translation models built on LLMs. You will combine expert and synthetic data to shape model behavior post‑training and drive production‑grade research with large parameter counts. You’ll own evaluation, reward modeling, and multi‑modal quality improvements, partnering with engineering to deploy at scale. This is a hands‑on, impact‑driven role within a collaborative research team at a mission‑centric AI company.
Leistungen / Benefits
- Hybrid work
- 30 days of annual leave
- Virtual Shares
- Hack Fridays
- Regular in‑person team events
- Mentally healthy resources
Verantwortungsbereiche
- Drive development of steerable translation models conditioned on user preferences, rules, and context
- Lead post‑training activities: supervised fine‑tuning, knowledge distillation, preference optimization, and RL tuned to translation quality
- Build reward and evaluator models for translation and investigate reward hacking and quality‑estimation failures
- Advance models ingesting multimodal content and context to boost translation quality
- Own full model delivery lifecycle: prototyping, ablations, training, evaluation, optimization, and production deployment
- Establish evaluation, reproducibility, monitoring, and continuous improvement practices in production
- Mentor researchers and engineers, promoting hands‑on collaboration and high model quality
Zentrale Anforderungen
- Proven experience making large models steerable and instruction‑following using methods like instruction tuning, latent space methods, steering vectors, or constrained encoding/decoding
- Deep hands‑on expertise in LLM post‑training (SFT, DPO), knowledge distillation, and/or RLHF/RLAIF, PPO/GSPO, and reward modeling
- Strong data‑centric instincts for synthetic data and preference data pipelines, LLM‑as‑judge generation, data curation and filtering
- Experience designing evaluation and reward signals using automatic metrics, LLM‑as‑judge, non‑verifiable rewards, and human‑in‑the‑loop evaluation
- Hands‑on experience training models, running experiments, debugging pipelines, and shipping ML systems to production with product impact
- Strong coding and experimentation skills (Python, PyTorch/JAX/TensorFlow) and clear communication to align with product and engineering priorities
- Collaborative mindset
- Mentoring and leadership
- Strong communication
- Python
- PyTorch
- JAX
Senior Research Scientist | Model Steering in Bonn Arbeitgeber: DeepL
DeepL ist ein hervorragender Arbeitgeber, der eine dynamische und unterstützende Arbeitsumgebung bietet, in der Innovation und persönliches Wachstum gefördert werden. Mit flexiblen Arbeitszeiten und der Möglichkeit, remote zu arbeiten, ermöglicht DeepL seinen Mitarbeitern, ihre Work-Life-Balance zu optimieren, während sie an bedeutungsvollen Projekten im Bereich KI-Technologie arbeiten. Die Unternehmenskultur basiert auf offener Kommunikation und Teamzusammenhalt, was durch regelmäßige Teamevents und monatliche Hackdays unterstützt wird, um Kreativität und Zusammenarbeit zu fördern.