- Experience
- 5–8 yrs
- Salary
- —
- Openings
- 1
- Posted
- 6 days ago
- Work mode
- In office
- Education
- Bachelor's degree
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About the Role
We are seeking a skilled Datadog Engineer to lead the management and enhancement of our observability platform. This role involves designing and implementing comprehensive monitoring solutions to ensure optimal system performance and reliability across cloud and containerized environments.
Key Responsibilities
- Oversee and fine-tune the Datadog observability platform and its monitoring setups.
- Create and maintain dashboards, alerts, service maps, and tailor-made monitoring tools.
- Deploy application performance monitoring (APM), distributed tracing, real user monitoring (RUM), log analytics, and synthetic monitoring capabilities.
- Establish and manage key performance indicators (KPIs), service level indicators (SLIs), and service level objectives (SLOs) to maintain service dependability.
- Integrate Datadog with incident management systems like ServiceNow, PagerDuty, and OpsGenie.
- Work closely with development, infrastructure, and Site Reliability Engineering (SRE) teams to diagnose and resolve performance bottlenecks.
- Automate monitoring workflows and generate reports using scripting languages such as Python, PowerShell, or Bash.
- Provide documentation on best practices and deliver expert guidance to team members on monitoring strategies.
Candidate Profile
- Possesses a Bachelor’s degree in Computer Science, Information Technology, or an equivalent field.
- Has 5 to 8 years of professional experience in monitoring, infrastructure management, or application support.
- Demonstrates over 3 years of direct experience working with Datadog’s APM, Logs, Metrics, RUM, and Synthetics features.
- Familiar with cloud platforms such as AWS, Azure, or Google Cloud Platform.
- Knowledgeable in container orchestration and development tools including Kubernetes, Docker, and CI/CD pipelines, alongside understanding distributed systems.
- Strong grasp of observability concepts involving metrics collection, log aggregation, and trace analysis.
- Experience with infrastructure as code using Terraform and understanding ITIL processes is a plus.
Minimum education
Bachelor's Degree
Skills
Tools & software
Docker
required
Kubernetes
required
Amazon Web Services AWS
required
Microsoft PowerShell
required
Google Cloud Platform
required
Microsoft Azure
required
Terraform
required