Senior MLOps Engineer (ML Workflows Engineering)
Berlin, Germany · Full Time
Be the first to apply
- Experience
- Any
- Salary
- —
- Openings
- 1
- Posted
- 15 ঘন্টা আগে
- Work mode
- In office
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About JetBrains and Our Mission
JetBrains has been passionate about developing the most powerful and efficient developer tools since 2000. Our goal is to automate routine tasks in coding, speeding production so that developers can focus on growth, creativity, and discovery.
Currently, AI-driven assistance and agents are becoming integral features in our IDEs. Our ML Workflows Engineering team is committed to eliminating infrastructure complexities, optimizing MLOps processes, and enabling teams to concentrate on creating impactful ML models and smart agents. As a member of this team, you will be pivotal in crafting tools, automations, and pipelines that make ML development more intuitive and streamlined.
By embracing advanced MLOps methodologies and engineering excellence, we seek to enhance productivity and simplify ML infrastructure to empower our teams to innovate boldly in AI.
Key Responsibilities
- Create tools, automations, and streamlined workflows that reduce infrastructure burdens and enable AI teams to prioritize experimentation and core challenges.
- Develop comprehensive monitoring, logging, and tracing solutions to guarantee reliable performance and reproducibility of ML workflows in production environments.
- Design, build, and maintain full-cycle machine learning pipelines that facilitate smooth development, training, and deployment of ML models and intelligent agents.
- Manage large-scale distributed systems, such as GPU clusters, to support ML model training, fine-tuning, and evaluation.
- Collaborate closely with product and development teams to convert high-level objectives into scalable, maintainable system designs and implementations.
- Optimize workflows for reproducibility, scalability, and cost-effectiveness while ensuring ML teams remain productive and innovative.
Candidate Requirements
- Demonstrated hands-on experience with contemporary MLOps tools including Kubernetes, cloud platforms like GCP and AWS, and ML orchestration frameworks.
- Strong understanding of the end-to-end ML lifecycle from ideation to deployment in customer-facing applications.
- Capability to independently manage projects from initial concept through design, experimentation, execution, and iteration.
- A user-focused approach, able to translate ML engineers' needs into scalable and maintainable architectural solutions.
- Familiarity with modern CI/CD tools such as GitHub Actions or JetBrains TeamCity.
- At least three years of experience using Python to write clean, maintainable code within modern ML codebases.
Preferred Additional Experience
- Working with ML orchestrators and workflow tools including ZenML, Dagster, and Airflow.
- Infrastructure development and service creation using Kubernetes and similar cluster solutions.
- Developing backend services with Python.
- Building and maintaining ML pipelines, including legacy systems.
- Utilizing experiment tracking and observability tools such as Weights & Biases, MLflow, Langfuse, or equivalents.
Highly Valued Expertise
- Experience with inference frameworks for large language models like vLLM, DeepSpeed, and TensorRT.
- Maintaining Python libraries used internally or externally by ML engineers.
- Strong theoretical knowledge in natural language processing and transformer architectures.
- Writing production code in Java and/or Kotlin.
Commitment to Diversity, Equity, and Inclusion
We are committed to providing equal opportunities and fostering an inclusive environment that welcomes individuals from all backgrounds, identities, religions, ages, accessibility needs, and orientations.
Privacy Notice
Your application data will be handled in accordance with our Recruitment Privacy Policy.