About us
Albs is an AI research lab building real-time multimodal intelligence for machines, enabling them to see, hear, reason, and interact. We treat model architecture, inference, and runtime as one system. Our purpose-built models and optimized runtimes unlock the full potential of each device within defined limits for hardware cost, power consumption, and response time. Companies can adapt our technology to their own machines without building the underlying AI from scratch.
We founded Albs at the intersection of LLM architecture research and on-device engineering, and we are currently in stealth, but well funded, with dedicated compute for large-scale training and experimentation. Publishing is a core part of our research culture, and we contribute our work to top venues. We share more details about the company, the team and our backing in the first conversation.
Who we're looking for
This is a Staff / Senior IC role. We are looking for experienced researchers and engineers, typically with a PhD and several years of research or industry experience or an equivalent track record. The exact scope of each role depends on your background: some people go deep on one part of the stack, others shape the technical direction of a whole area. We agree on scope together with you during the interview process. We welcome applications from all qualified candidates, regardless of gender, age, ethnic origin, religion, disability or sexual orientation.
What you'll work on
-
Develop and train our vision-language models end to end, from vision encoders and training targets to data pipelines for rich, generalizing representations.
-
Build real-time visual understanding for machines: streaming video, grounding and spatial reasoning at low latency.
-
Design for the hardware from the start: research efficient architectures, visual token reduction and resolution strategies for on-device inference, together with the hardware optimization team.
-
Build evaluations for grounding, video understanding and real-world robustness.
-
Explore ways to close the gap between VLMs and world models, including joint training of LLMs and VLMs in new architectures.
-
Integrate the vision component into our multimodal language model without regressing text performance.
-
Publish at top venues and help shape our research agenda in vision-language models.
What we're looking for
-
You hold a PhD in a relevant field and have several years of research experience after it, in academia or industry, or an equivalent track record.
-
You have pre-trained or significantly advanced vision-language models end to end.
-
You have a deep understanding of how vision and language representations interact: tokenization, alignment, grounding, cross-modal attention, and the failure modes of each.
-
You have a strong publication record (e.g. CVPR, ICCV, ECCV, NeurIPS, ICML, ICLR) or an unambiguous production track record in multimodal models.
-
You have deep Python and PyTorch proficiency and experience with distributed training at multi-node scale.
-
You are as comfortable in the training code as in the literature, and you test your ideas against real hardware constraints.
-
Bonus: experience with video models, efficient VLMs, or omni-modal training.
How We Work Together
We are a small, focused team of experts. You would join early, work directly with the founders, and help shape how we build. Fast iteration, short lines of communication, and in-person discussion matter a lot to us. Our culture is built around the office in Freiburg, Germany, with a default of three days a week on site. Alternatively, you can work remotely and join us on a regular cadence. We will discuss what works best for you during the interview process.
We hold ourselves to a high standard: the research must be rigorous, the understanding deep, and the product well crafted. Ideas are judged on their merit, not on who proposed them, credit belongs to the team, and no task is beneath anyone. We make ambitious bets and ship early instead of waiting for perfect conditions. Above all, we care about each other, because initiative only works when it comes with respect for the people around you.
What we offer
Expect a competitive base salary plus equity, and dedicated compute for your research and experiments. We give you the time and support to publish and present at top conferences, and flexibility in how you work, whether on site or remote with regular in-person visits. With flat hierarchies and short decision paths, good ideas move from discussion to experiment quickly. And yes, the coffee is excellent, and we regularly get together as a team outside of work. We go through compensation and all other details with you early in the process.
#J-18808-Ljbffr
Member of Technical Staff - Vision-Language Models in Freiburg im Breisgau Arbeitgeber: Albs Labs GmbH
Albs ist ein hervorragender Arbeitgeber, der eine dynamische und unterstützende Arbeitsumgebung in Freiburg, Deutschland, bietet. Mit einem starken Fokus auf Forschung und Entwicklung im Bereich KI ermöglicht das Unternehmen seinen Mitarbeitern, an bedeutenden Projekten zu arbeiten und ihre Ideen schnell umzusetzen. Die flachen Hierarchien und die Flexibilität, sowohl vor Ort als auch remote zu arbeiten, fördern das persönliche Wachstum und die Zusammenarbeit im Team.