- Experience
- 3+ yrs
- Salary
- —
- Openings
- 1
- Posted
- 1 ದಿನ ಹಿಂದೆ
- Work mode
- In office
- Education
- Bachelor's degree
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About the Role
We are looking for an experienced Databricks Data Engineer to become a key member of our data platform team. This position involves designing, constructing, and optimizing both batch and real-time data pipelines using the Databricks Lakehouse platform. You will be responsible for applying best practices with technologies such as Delta Lake, Apache Spark, and Unity Catalog to build data infrastructure that is scalable, secure, and cost-efficient.
Responsibilities
- Design and deploy strong data pipelines leveraging Apache Spark, PySpark, and Spark SQL to gather data from multiple sources, including APIs, relational databases, cloud storage systems, and event streams.
- Implement and maintain the Medallion Architecture layers (Bronze, Silver, Gold) with Delta Lake to provide clean, transactional, and analytics-ready datasets.
- Automate workflows for batch and real-time streaming using Databricks Lakeflow Jobs, Workflows, or Apache Airflow.
- Configure granular access controls, security protocols, and data lineage using Unity Catalog.
- Enhance performance by tuning Spark jobs, optimizing query execution, and managing cluster configurations such as autoscaling, caching, and data partitioning to balance performance and cloud cost efficiency.
- Collaborate with Data Scientists, Business Intelligence Engineers, and product teams to convert business needs into maintainable data models. Implement unit testing and CI/CD for data engineering deliverables.
Required Qualifications
- Bachelor's degree in Computer Science, Data Engineering, Information Systems, or equivalent experience.
- Minimum of 3 years hands-on experience in data engineering, including at least 2 years building solutions on the Databricks platform.
- Strong skills in Python (PySpark) and SQL; Scala knowledge is advantageous.
- Proficiency with Apache Spark and Delta Lake core technologies.
- Experience working with cloud platforms such as AWS, Azure, or Google Cloud Platform.
- Knowledge of Unity Catalog for data governance and Databricks Workflows for orchestration.
- Familiarity with Git version control, CI/CD tools, and writing modular, testable code.
Preferred Qualifications
- Databricks Certified Data Engineer Associate or Professional certification.
- Experience with real-time streaming systems like Structured Streaming, Apache Kafka, or Kinesis.
- Understanding of dbt (data build tool) and its integration with Databricks.
- Exposure to machine learning operations, notably MLflow, and vector databases within Databricks.
Minimum education
Bachelor's Degree
Skills
Tools & software
Git
required
Apache Spark
required
Apache Airflow
required
Databricks
required