- Experience
- 5+ yrs
- Salary
- —
- Openings
- 1
- Posted
- 1 week ago
- Work mode
- In office
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
Job Overview
We are seeking an AI Engineer who will be responsible for developing, deploying, and managing AI and large language models (LLM) within a dual-environment infrastructure combining Google Cloud Platform (GCP) for public cloud workloads and a sovereign cloud (Humain) for handling classified data.
Key Responsibilities
- Create and fine-tune machine learning and LLM models focused on Arabic natural language processing, document classification, computer vision and optical character recognition (OCR), as well as artificial intelligence operations (AIOps) use cases.
- Conduct thorough pre-deployment evaluations including setting accuracy baselines, regression analysis, safety testing, and provide justifications for GPU resource allocations.
- Optimize model inference processes through techniques such as quantization, batching, and context sizing based on actual usage metrics.
- Deploy services in the sovereign cloud environment using Humain GPU-as-a-Service, employing Kubernetes with GPU partitioning on B300 nodes along with quota management and role-based access control (RBAC).
- Replicate workloads on the GCP platform utilizing Vertex AI and Google Kubernetes Engine (GKE), implementing classification-based routing strategies.
- Maintain ownership over the AI model serving stack including vLLM/TGI frameworks, model version control, continuous integration and delivery (CI/CD), and monitoring performance indicators such as latency, token consumption, GPU utilization, and concept drift.
- Ensure all AI models comply with Saudi Arabia's Zakat, Tax and Customs Authority (ZATCA) data sovereignty policies and the Saudi Data and AI Authority (SDAIA) regulations, including AI Ethics, Generative AI guidelines, and Personal Data Protection Law (PDPL).
Requirements
- At least five years of professional experience in machine learning and AI engineering with a proven track record in deploying production-ready LLM solutions.
- Strong programming proficiency in Python and experience with PyTorch and Hugging Face libraries.
- Hands-on experience deploying Kubernetes in production environments with GPU-based inference capabilities.
- Expertise in cloud platforms especially Google Cloud Platform’s Vertex AI or comparable cloud services.
Additional Information
The role is based in Riyadh, Saudi Arabia and requires onsite commitment.