PhysicsX

Senior Machine Learning Infrastructure Engineer, Research

PhysicsX

Singapore · Full Time

Be the first to apply

Experience
5+ yrs
Salary
Openings
1
Posted
6 giorni fa
Work mode
In office
Resume
Required to apply

Where you'll work

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

About PhysicsX

PhysicsX is an advanced technology company rooted in numerical physics and Formula One heritage, driving hardware innovations with the speed of software development. The team develops AI-driven simulation software for engineering and manufacturing across high-tech sectors like Aerospace, Defense, Materials, Energy, Semiconductors, and Automotive. This software enables precise, multi-physics simulation powered by AI inference, helping engineers enhance optimization, automation, and innovation throughout the entire engineering lifecycle.

Role Overview

The Senior Machine Learning Infrastructure Engineer will be responsible for developing and managing the infrastructure supporting research operations related to model training, fine-tuning, and deployment pipelines. The position is tightly integrated within the Research team, collaborating closely with both ML engineers and research scientists to support efficient and scalable training of Large Physics Models.

Key Responsibilities

  • Create and maintain distributed neural network training systems for advanced architectures such as Transolver and Point Cloud Transformer using large-scale NVIDIA DGX B200 clusters.
  • Enhance training workflows focusing on improved throughput, robustness, and cost-effectiveness, including checkpointing strategies, gradient accumulation, and synchronization across multiple nodes.
  • Develop and uphold experiment monitoring and observability platforms to provide clear insights into training metrics, hyperparameter tuning, and model outputs.
  • Address data input/output challenges related to handling extensive mesh datasets, optimizing cloud-storage data pipelines through prefetching, caching, and data formatting improvements.
  • Design and implement reliable model serving frameworks for pre-trained Large Physics Models with features such as zero-shot inference and uncertainty quantification via Monte Carlo Dropout.
  • Build model packaging and deployment processes ensuring dependable execution and fine-tuning abilities in customer environments while guaranteeing reproducibility of results.
  • Improve developer workflows to enable rapid iterations, dependable continuous integration/deployment, and effective debugging tools within the Research team.
  • Work collaboratively with the broader infrastructure group to establish company-wide infrastructure standards and shared best practices.

Candidate Qualifications

  • Proven ability to scope, prioritize, and deliver complex infrastructure projects efficiently.
  • Strong analytical abilities for diagnosing infrastructure issues and devising solutions promptly.
  • Exceptional communication skills capable of translating research challenges into infrastructure strategies and implementations.
  • Minimum of 5 years' experience in building and operating large-scale machine learning infrastructure.
  • In-depth knowledge of distributed training systems including debugging NCCL issues, optimizing communication, and balancing approaches such as Fully Sharded Data Parallel (FSDP), Distributed Data Parallel (DDP), and pipeline parallelism.
  • Solid foundation in systems engineering: Linux OS, advanced networking (NVLink, InfiniBand), storage I/O, performance profiling and optimization.
  • Hands-on experience with Kubernetes and SLURM for orchestrating GPU cluster workloads.
  • Advanced proficiency in Python programming and ML frameworks, particularly PyTorch.
  • Experience managing cloud GPU environments, preferably CoreWeave or similar platforms optimized for GPU/HPC usage.

Preferred Experience

  • Familiarity with geometric deep learning techniques and neural operators working with mesh, point cloud, or graph data.
  • Background in high-performance computing for simulation engineering, especially workflows involving Computational Fluid Dynamics (CFD) or Finite Element Analysis (FEA).
  • Expertise in constructing model serving systems meeting latency and throughput requirements.
  • Knowledge of experiment tracking tools such as Weights & Biases or MLflow and monitoring tools like Prometheus and Grafana.
  • Experience packaging models for deployment including containerization, model registries, and version control.

Benefits and Culture

  • Opportunity to develop impactful AI-native engineering technologies at an influential stage with real-world industrial applications.
  • Collaborate with a diverse and talented group of engineers, scientists, and operators dedicated to excellence and mutual growth.
  • Flat organizational structure promoting idea meritocracy, openness to questioning and innovating.
  • Balance between purposeful work and a sustainable lifestyle with a hybrid work model combining on-site and remote work.
  • Equity participation in the company.
  • 10% employer pension contribution supporting future financial security.
  • Free office meals to keep energy levels high.
  • Generous parental leave with 3 months full pay for paternity and 6 months full pay for maternity.
  • Childcare assistance via the YellowNest nursery scheme.
  • 25 days annual leave plus public holidays to ensure adequate rest.
  • Comprehensive private medical insurance fully covered for employees.
  • Wellness program subscription offering gym memberships, classes, and wellness apps.
  • Regular eye examinations.
  • Support for career development and learning opportunities.
  • Confidential employee assistance program for well-being.
  • Green commuting incentives including Bike2Work, season ticket loans, and electric vehicle salary sacrifice schemes.

Diversity and Inclusion

The company is dedicated to fostering a diverse workforce and providing equal employment chances regardless of gender, ethnicity, disability, age, sexual orientation, or other identities. It actively encourages candidates from underrepresented groups in technology fields to apply and supports women from disadvantaged backgrounds through educational sponsorships. Data on diversity is collected strictly for policy monitoring and remains confidential.

Level

Senior

Tools & software

PyTorch PyTorch required

How they work

Communication Teamwork & Collaboration Problem Solving

Leave it if you'd like a reply — we won't use it for anything else.

Click to browse, drag & drop, or paste a screenshot

PNG, JPG, GIF, MP4, WebM, MOV · Max 20MB each · Up to 5 files

🤖
Online · instant AI help
Broxer