- Experience
- 3+ yrs
- Salary
- —
- Openings
- 1
- Posted
- 4 weeks ago
- Work mode
- In office
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
Role Overview
We are seeking a proficient Scala and PySpark Developer with practical experience in Hadoop ecosystems and expertise in CI/CD GitOps methodologies. The successful candidate will be responsible for designing, building, and tuning large-scale data processing systems while enhancing automated deployment pipelines.
Core Responsibilities
- Create, develop, and maintain scalable data pipelines and batch processing frameworks using Scala and PySpark.
- Utilize Hadoop components extensively to handle distributed data storage and processing.
- Enhance performance, scalability, and reliability of Spark jobs working on vast datasets.
- Build and manage CI/CD pipelines adhering to GitOps principles to ensure automated deployment and stable environments.
- Work closely with data engineering, architecture, DevOps teams, and business users to deliver solid data solutions.
- Conduct code reviews, write unit tests, debug issues, and provide production support for data-related applications.
- Maintain compliance with coding standards, security policies, and operational excellence.
Required Qualifications
- Hands-on experience in Scala and PySpark programming.
- Proven experience with Hadoop and other distributed data processing technologies.
- Practical knowledge of CI/CD tools and deployment based on GitOps workflows.
- Understanding of data engineering principles including ETL/ELT pipelines and big data system architectures.
- Experience in diagnostics, performance tuning, and application support in production environments.
- Strong analytical thinking, communication capabilities, and teamwork skills.
Desirable Skills
- Familiarity with Kafka and event-driven or data streaming systems.
- Experience with monitoring and observability platforms like AKHQ, Prometheus, and Grafana.
- Knowledge of workflow orchestration tools such as Apache Airflow.
- Exposure to cloud computing environments and container orchestration technologies is advantageous.
Experience Requirements
Typically requires a minimum of three years’ experience in big data engineering or platform development.
Tools & software
How they work
Communication
Teamwork & Collaboration
Problem Solving
Attention to Detail