- Experience
- 3+ yrs
- Salary
- —
- Openings
- 1
- Posted
- 2 weeks ago
- Work mode
- Work from home
- Resume
- Required to apply
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About Social Discovery Group
Social Discovery Group (SDG) comprises a collection of companies focused on social discovery solutions, aiming to address challenges like loneliness, isolation, and disconnection by redefining virtual social interactions. Their platforms connect people worldwide across diverse cultures and regions.
The company employs a remote international team of IT specialists and digital nomads committed to creating impactful social discovery products. SDG has earned recognition as a two-time "Great Place to Work" in the USA and Japan (2024–2025) and is ranked among the Top 5 companies for work-from-anywhere jobs by FlexJobs (2025).
Role Overview
We are seeking a skilled Site Reliability Engineer (SRE) enthusiastic about ensuring infrastructure reliability, automation and building scalable production environments.
Key Responsibilities
- Manage and enhance the reliability and stability of production infrastructure.
- Plan, carry out, and support infrastructure deployments and modifications.
- Develop and maintain Infrastructure-as-Code solutions with tools like Ansible and Terraform.
- Administer and optimize containerized and Kubernetes-based environments.
- Create automation scripts and internal tools to streamline operations.
- Monitor system health, analyze incidents and proactively enhance observability.
- Collaborate with Development, QA, DevOps, and other SRE teams to improve CI/CD processes.
- Manage monitoring and alerting systems to minimize downtime and boost system performance.
- Document technical processes thoroughly with runbooks and operational guides.
- Provide support for DNS, Web Application Firewall (WAF), CDN, and caching infrastructures as needed.
Candidate Requirements
- Minimum 3 years’ experience in SRE, DevOps, System Administration, or Build/Release Engineering roles.
- Proficient in Linux system administration and troubleshooting.
- Hands-on experience with Kubernetes and container technologies like Docker or Podman.
- Familiarity with CI/CD pipelines, preferably using GitLab CI.
- Solid knowledge of Infrastructure-as-Code and configuration management, especially Ansible and/or Terraform.
- Experience working with monitoring and observability tools such as Prometheus, Grafana, Zabbix, or VictoriaMetrics.
- Strong understanding of networking concepts including DNS, HTTP/HTTPS, load balancing, and diagnosing network issues.
- Proficient understanding of Git and contemporary software delivery workflows.
- Ability to work autonomously, demonstrate ownership, and proactively enhance infrastructure.
- Fluent in Russian at a technical documentation and communication level.
Preferred Additional Skills
- Experience with cloud platforms like AWS and Google Cloud Platform.
- Knowledge of RabbitMQ or AMQP protocols.
- Familiarity with Cloudflare, Akamai, WAF, and CDN services.
- Exposure to tracing and advanced observability tools.
Benefits and Perks
- Fully remote, full-time opportunity.
- Annual vacation entitlement of 28 calendar days.
- 7 wellness days annually for personal matters or recuperation without using sick leave.
- Referral bonuses up to $5,000 for successful candidate recommendations.
- 50% funding support for professional training, international conferences, and meetings.
- Corporate discounts on English language courses.
- Health benefits including up to $1,000 gross yearly compensation toward private health insurance or medical expenses for the employee and close family members if corporate insurance is unavailable.
- Workspace setup provision including equipment or reimbursement up to $1,000 every 3 years for home office or co-working spaces.
- Internal reward system with bonuses redeemable for merchandise, team events, massages, and more.