Applied Scientist, Agent Evaluation & Adaptive Model Routing
Singapore · Full Time
Be the first to apply
- Experience
- Any
- Salary
- —
- Openings
- 1
- Posted
- منذ 5 أيام
- Work mode
- In office
- Education
- Bachelor’s or higher in relevant technical discipline
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About Bitdeer
Bitdeer is a globally recognized technology leader specializing in Bitcoin mining and AI cloud computing. The company offers end-to-end Bitcoin mining solutions, including ASIC chip design, mining rig manufacturing, and overseeing operations related to equipment procurement, logistics, data center design, and network management. Bitdeer operates worldwide with a 3 GW energy portfolio, managing mining and HPC data centers in various countries such as the US, Bhutan, Norway, Canada, Malaysia, and Ethiopia. Headquartered in Singapore, Bitdeer emphasizes innovation in computing power solutions.
About Bitdeer AI Lab
As a pioneering AI research lab under Bitdeer, the AI Lab focuses on advancing artificial intelligence technologies with a mission to make inference affordable and efficient by optimizing the power, data centers, and software infrastructure supporting AI applications.
Role Responsibilities
- Develop and enhance systems for evaluating and making agentic inference reliable, measurable, and adaptive.
- Own and improve the LLM (large language model) and agent evaluation pipelines, including creating representative task suites for evaluation.
- Analyze agent behavior on multiple metrics such as task success, tool usage, quality, cost, latency, and token consumption at the trajectory level.
- Research and prototype adaptive model routing strategies within a shared agent framework, including rule-based systems, cascades, stage-aware routing, uncertainty-informed selection, and escalation/recovery policies.
- Advance promising methods from offline evaluations through traffic replay, shadow testing, and pilot projects.
- Collaborate closely with MaaS and platform engineering teams to integrate solutions into production, sharing responsibilities for gateway reliability, billing, access control, and service level agreements.
Required Qualifications and Experience
- Degree (Bachelor’s, Master’s, or PhD) in Computer Science, Machine Learning, Statistics, Electrical Engineering, or related disciplines.
- Extensive practical experience in large language model evaluation, agent systems, applied machine learning, or adaptive inference methodologies.
- Proficiency in Python programming and familiarity with PyTorch and contemporary data/evaluation tools to independently create robust experimental pipelines and research prototypes.
- Experience designing or managing evaluation pipelines for LLMs or agents, covering task and trajectory-level metrics, dataset construction, automated scoring, regression testing, and failure analysis.
- In-depth knowledge of at least one core specialty such as model selection and routing, uncertainty estimation/calibration, cascading and escalation strategies, or stage-aware agent inference.
- Bonus skills include experience with preference modeling, contextual bandits, online learning, or other adaptive inference techniques.
- Hands-on evaluation of multi-turn or tool-using agents, with attention to task completion, tool-call accuracy, planning and recovery behaviors, and balanced cost/latency considerations.
- Strong experimental methodology including controlled comparisons, honest baselines, and clear interpretation of evaluation outcomes.
- Ability to analyze performance per request and trajectory, constructing cost-quality trade-off frontiers beyond aggregate scores.
- Preferred experience with multi-model APIs, agent harnesses, traffic replay, shadow testing, A/B testing, or production environment model monitoring.
- Additional familiarity with model-specific differences in tool usage, context window management, prompt caching, reasoning controls, and inference frameworks like vLLM or SGLang is advantageous.
- Contributions to top-tier ML, NLP, or systems research publications or significant open-source projects in evaluation, agents, routing, or inference are welcomed.
- Demonstrated initiative and product sense, with experience advancing ambiguous research questions to working prototypes and measurable validation.
Work Environment and Benefits
- Inclusive workplace that values diversity and authentic expression.
- Open office layout fostering collaboration and a startup-inspired atmosphere.
- Opportunity to connect with industry leaders and enthusiasts within a rapidly growing company.
- Direct impact on the evolution of the digital asset sector through meaningful contributions.
- Engagement in innovative projects and development of operational processes and systems.
- Personal growth supported by responsibility, autonomy, and learning opportunities.
- Competitive benefits package including welfare programs, training, and mentorship.
Minimum education
Bachelor's Degree