- Experience
- 30+ yrs
- Salary
- CAD 30 – CAD 50 / hour
- Openings
- 1
- Posted
- 2 weeks ago
- Work mode
- Work from home
- Resume
- Required to apply
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
Overview
This position involves designing, testing, and refining prompts to guide large language models (LLMs) in producing accurate, safe, and valuable outputs within various production workflows. You will work closely with engineering teams and AI/ML data operations to enhance model behavior through iterative prompt development, evaluations, and feedback processes related to reinforcement learning from human feedback (RLHF).
Key Responsibilities
- Craft system, user, and tool prompts for LLM-based products such as chatbots, search engines, summarization tools, extraction utilities, and coding assistants.
- Conduct prompt and quality assurance evaluations using defined rubrics assessing accuracy, grounding, safety, and style compliance.
- Develop and maintain datasets for prompt testing including labeling and preference data that align with RLHF goals.
- Investigate prompt failures such as hallucinations, instructional gaps, or unsafe outputs, and recommend modifications to prompts, policies, or data.
- Create and enforce annotation guidelines, ensuring the quality and compliance of training datasets.
- Support training pipelines for LLMs including regression testing and efforts to boost model performance.
- Participate in safety content labeling, refusal behavior checks, and defenses against jailbreaking or misuse.
- Maintain documentation for prompt libraries, evaluation datasets, and decision logs to ensure reproducibility in experiments.
Required Qualifications
- Mid-level to senior experience in developing or optimizing prompts for LLM applications in live or rigorous testing contexts.
- Exceptional writing and analytical capabilities to clearly translate product goals into precise prompts and restrictions.
- Practical experience with rubric-driven prompt and QA evaluations.
- Understanding of RLHF principles including preference data and balancing helpfulness versus harmlessness.
- Capability to establish measurable standards for prompt quality and safety even when requirements are ambiguous.
Preferred Qualifications
- Experience with natural language processing tasks such as summarization, classification, named entity recognition, extraction, and retrieval-augmented generation (RAG) evaluation.
- Knowledge of content safety protocols, policy formulation, and security testing including red-teaming.
- Experience in constructing evaluation data sets and managing quality assurance initiatives.
- Basic scripting and data analysis skills using spreadsheets, SQL, or Python to support tracking and troubleshooting.
Compensation and Work Setup
The hourly wage ranges from 30 to 50 Canadian dollars. The role is fully remote and available across all Canadian time zones.