Forward Deployed Engineer - Full Stack, Model Optimization & AI Fine-Tuning
Chennai, Tamil Nadu, India · Full Time
Be the first to apply
- Experience
- 8–15 yrs
- Salary
- —
- Openings
- 1
- Posted
- 12 minutes ago
- Work mode
- In office
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
Overview
This role is geared toward an elite Applied AI Engineer who excels beyond routine API utilization and dives deep into model-level optimization. The candidate must possess a thorough understanding of attention mechanism mathematics, GPU performance tuning, and domain-specific model customizations. They will be responsible for high-level technical vision and managing complex edge cases.
Primary Responsibilities
- Fine-tune open-source large models like Llama 3 and Mistral using techniques such as Parameter-Efficient Fine-Tuning (PEFT), LoRA, and QLoRA to tailor models for client-specific domains.
- Optimize and quantize models to reduce inference costs and latency while maintaining output quality, including management of Dense Vectors and embedding optimizations.
- Engage in continuous research to integrate state-of-the-art advancements such as State Space Models and long-context optimizations into client solutions.
- Serve as a strategic consultant to C-level client executives by shaping the "Art of the Possible" and advising on long-term AI technology roadmaps.
Required Technical Expertise
- Strong proficiency in deep learning frameworks like PyTorch and TensorFlow, with deep understanding of Transformers architecture internals and attention mechanisms.
- Experience in Model Operations including serving custom models (e.g., vLLM, TGI), advanced GPU memory management, and quantization methods such as GGUF and AWQ.
- Expertise in advanced data handling, including training data curation, synthetic data generation, and reinforcement learning with human feedback (RLHF) concepts.
- Leadership capabilities to establish technical culture and set standards across the Forward Deployed Engineering organization.
Tools & software
PyTorch
required
TensorFlow
required