- Experience
- 6+ yrs
- Salary
- —
- Openings
- 1
- Posted
- 2 days ago
- Work mode
- In office
- Resume
- Required to apply
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About the Role
This position entails full ownership of the machine learning model lifecycle in production environments, ensuring that validated models are effectively deployed, continuously monitored, optimized, and maintained at scale. The focus areas include MLOps practices, cloud platform management, automation, observability enhancements, cost efficiency, and handling incidents during production.
Key Responsibilities
- Manage the serving, continuous integration and deployment (CI/CD), deployment processes, and ongoing monitoring of ML models in production.
- Develop and sustain pipelines for model retraining, version control, and deployment.
- Oversee ML infrastructure while applying financial operations (FinOps) methods to optimize cloud resource expenditure.
- Implement comprehensive observability measures, including alert systems, detection of model performance drift, and continuous performance tracking.
- Take ownership of production incident management including response and troubleshooting efforts; participate in on-call duties as required.
- Collaborate closely with data scientists and ML engineers to create production-ready, scalable ML systems.
- Assist presales teams and proof-of-concept demonstrations by showcasing the scalability and production readiness of ML solutions.
- Provide guidance and mentorship to engineering team members on best practices for MLOps and maintaining production readiness.
- Effectively communicate trade-offs among infrastructure costs, system performance, and reliability to non-technical stakeholders.
Candidate Qualifications
- A minimum of six years of professional experience in MLOps, ML platform engineering, or related fields with demonstrable ownership of production deployments.
- Proficient in CI/CD methodologies, containerization technologies, cloud platforms, and ML observability frameworks.
- In-depth knowledge of the machine learning model lifecycle, including retraining workflows, versioning strategies, and model drift detection techniques.
- Experience utilizing Infrastructure as Code (IaC) and developing automated deployment pipelines.
- Strong familiarity with leading cloud service providers such as AWS, Microsoft Azure, or Google Cloud Platform (GCP).
- Hands-on experience in cloud cost monitoring, optimization initiatives, and implementing FinOps strategies.
- Comprehensive understanding of monitoring systems, alert protocols, service level agreement (SLA) management, and managing incidents during production.
- Skilled in troubleshooting technical issues and effectively communicating incidents to business stakeholders.
- Excellent collaboration abilities and mentoring experience within engineering teams.
- Willingness and readiness to engage in on-call rotations and provide off-hour support for production environments.
Skills
How they work
Communication
Teamwork & Collaboration
Problem Solving
Leadership
Dependability