Researcher / PhD Student, Mechanistic Interpretability for Safe Agentic AI in Saarbrücken

Researcher / PhD Student, Mechanistic Interpretability for Safe Agentic AI in Saarbrücken

Saarbrücken Vollzeit Kein Homeoffice möglich
D

We seek a PhD researcher at the earliest possible date to advance post-hoc controllability of language models through mechanistic interpretability. You will develop methods for fine-grained activation steering that enable safe behavior modification without retraining, a critical requirement for deploying agentic systems in safety-critical domains.

The position is embedded in a collaborative research project combining mechanistic circuit analysis with safe AI system design. Your work will directly contribute to understanding how models can be reliably steered at the neuron and sparse feature level, while preserving previously learned capabilities.

We value rigorous, methodologically grounded research: you'll be expected to design careful experiments, publish in top venues, and engage in the kind of open, critical discussion of ideas that moves the field forward. This is basic research with real-world implications, grounded in our lab's commitment to explainable and efficient language processing.

Importantly, you will be part of a fully integrated interdisciplinary team spanning both SLIMS and SAgA projects. Your mechanistic findings will directly inform safety architectures, and you'll collaborate closely with researchers across interpretability, multilingual NLP, formal ethics, and agentic systems.

The position is embedded in the Multilinguality and Language Technology (MLT) group under the direction of Dr. Marius Mosbach, with broader departmental support from Prof. Kristian Kersting and Prof. Verena Wolf; the primary supervisor of the position is Dr. Simon Ostermann; the candidate will be based in his ‘Efficient and Explainable’ NLP team.


  • Develop and validate interpretable steering methods targeting specific computational circuits
  • Analyze failure modes and robustness of activation-level interventions under distribution shift and compositional scenarios
  • Design experiments quantifying orthogonality between safety interventions and task-critical representations
  • Contribute to scaling steering methods from proof-of-concept to production-grade robustness
  • Publish results in top-tier venues (NeurIPS, ICML, ACL, ICLR)


Required Background:

  • Master's degree or equivalent in computer science, NLP, machine learning, or related field
  • Strong foundation in deep learning and neural network architectures
  • Experience with mechanistic interpretability methods (circuit analysis, SAEs, activation patching, or similar)
  • Proficiency in Python and PyTorch or equivalent frameworks
  • Ability to work independently and in collaborative teams

Desirable Qualifications:

  • Prior experience with model steering, causal intervention, or mechanistic analysis
  • Familiarity with interpretability tooling (e.g., TransformerLens, Anthropic's SAE library)
  • Track record of publications or strong problem-solving demonstrated in prior work
  • Interest in AI safety and alignment research


  • We offer competitive, market-based compensation and many other benefits (Urban Sports Club, corporate benefits, a subsidy for your Jobticket, and much more)
  • Access to GPU clusters and computational resources
  • A supportive, intellectually rigorous research environment where critical discussion and methodology matter
  • Space to develop your own scientific identity and pursue ideas with genuine independence
  • Active publication culture with strong support for conference attendance and dissemination
  • Integration into the broader mechanistic interpretability research community
  • Collaborative, respectful lab culture that values both scientific rigor and human well-being

#J-18808-Ljbffr

Researcher / PhD Student, Mechanistic Interpretability for Safe Agentic AI in Saarbrücken Arbeitgeber: Deutsches Forschungszentrum für Künstliche Intelligenz GmbH

Das DFKI Saarbrücken bietet eine herausragende Arbeitsumgebung für Forscher im Bereich der erklärbaren KI, insbesondere in einem innovativen Projekt wie Q-Fin. Hier profitieren Sie von einer engen Zusammenarbeit mit führenden Institutionen wie dem Forschungszentrum Jülich und Deka Investment GmbH, während Sie gleichzeitig die Möglichkeit haben, einen Doktortitel an der Universität des Saarlandes zu erwerben. Unsere Unternehmenskultur fördert wissenschaftliche Neugier, interdisziplinäres Denken und bietet Zugang zu realen Finanzdaten, was Ihre berufliche Entwicklung in einem dynamischen und zukunftsorientierten Umfeld unterstützt.

D

Kontaktdaten:

Deutsches Forschungszentrum für Künstliche Intelligenz GmbH Recruiting-Team