Reinforcement Learning Research Engineer - Cybersecurity AI Startup
London, England, United Kingdom · Full Time
Be the first to apply
- Experience
- Any
- Salary
- GBP 350,000 / year
- Openings
- 1
- Posted
- 3 weeks ago
- Work mode
- In office
- Education
- PhD or equivalent experience
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About the Role
We are seeking an experienced Reinforcement Learning (RL) Research Engineer to lead the post-training development of advanced language models designed for real-time security threat detection. This role involves architecting and deploying RL pipelines—such as RLHF, DPO, and GRPO—to effectively classify malicious behavior patterns. Your contributions will be vital in building resilient autonomous AI systems, safeguarding enterprise security infrastructure.
Company Overview
This opportunity is with a well-capitalized cybersecurity AI startup based in London, focused on creating next-generation defense solutions for autonomous AI in an increasingly complex threat landscape. Supported by top venture capital firms, this seed-stage company offers a unique chance to impact security technology at a foundational level alongside a highly skilled founding team.
Key Responsibilities
- Design and implement RL post-training pipelines, including reinforcement learning with human feedback (RLHF), direct preference optimization (DPO), and generalized ranked preference optimization (GRPO), to enhance the performance of models in security threat classification.
- Develop and manage scalable, end-to-end training systems encompassing distributed training frameworks, evaluation tools, and robust regression testing processes.
- Create custom reward functions and optimize preference models aimed at detecting intricate and long-duration malicious agent behaviors.
Candidate Profile
- Demonstrated practical experience with RL techniques applied to language models, especially in post-training scenarios like RLHF or DPO, either in research or production environments.
- Strong engineering skills with proficiency operating complex machine learning training pipelines at scale, utilizing frameworks such as PyTorch, JAX, and distributed computing technologies.
- Academic background in machine learning research, preferably with a PhD or equivalent expertise focused on optimization, alignment, and reinforcement learning methodologies.
Additional Information
Recruitment for this position is handled by an AI partnership system that matches candidates from a network to ensure the best fit for both parties. While some client details remain confidential, all roles are verified and genuine. This position offers a competitive salary up to £350,000 including equity. The role is located onsite in London, UK.
Minimum education
Doctorate