Design, build, and optimize scalable Databricks-based ETL/ELT pipelines (batch and streaming) using PySpark, Spark SQL, and Delta Lake. Implement Medallion architecture, integrate diverse data sources, ensure data quality/security/governance, automate deployments via CI/CD, monitor production pipelines, and collaborate with data and business stakeholders.
Role: Databricks Engineer
Experience
Core Technologies
Mandatory Skills
Experience
4–8 years
As per business requirement
Full-time
We are looking for a skilled Databricks Engineer with strong expertise in designing, developing, and optimizing modern data engineering solutions on the Databricks Lakehouse Platform. The ideal candidate should have experience building scalable ETL/ELT pipelines, working with large-scale data, and leveraging Apache Spark to deliver high-performance data solutions.
- Design, develop, and maintain scalable data pipelines using Databricks.
- Build ETL/ELT workflows for batch and streaming data processing.
- Develop solutions using PySpark, Spark SQL, and Delta Lake.
- Implement Medallion Architecture (Bronze, Silver, Gold) for data transformation.
- Integrate data from various sources including relational databases, APIs, cloud storage, and streaming platforms.
- Optimize Spark jobs for performance, scalability, and cost efficiency.
- Collaborate with Data Architects, Data Scientists, BI developers, and business stakeholders.
- Implement CI/CD pipelines and deployment automation for Databricks workloads.
- Ensure data quality, security, governance, and compliance.
- Monitor, troubleshoot, and optimize production data pipelines.
- Document technical solutions and follow engineering best practices.
Core Technologies
- Databricks Lakehouse Platform
- Apache Spark
- PySpark
- Spark SQL
- Delta Lake
- Python
- SQL
- Microsoft Azure (preferred)
- AWS
- Google Cloud Platform
- Azure Data Factory (ADF)
- Azure Data Lake Storage (ADLS Gen2)
- Azure Synapse Analytics
- Azure Key Vault
- Azure DevOps
- Data Warehousing
- Data Modeling
- ETL/ELT Development
- Batch Processing
- Streaming (Kafka/Event Hubs)
- Data Lake Architecture
- Git
- Azure DevOps / GitHub
- CI/CD Pipelines
- Experience with Unity Catalog.
- Knowledge of Databricks Workflows and Jobs.
- Hands-on experience with Delta Live Tables (DLT).
- Exposure to MLflow is an added advantage.
- Experience with data governance and security best practices.
- Familiarity with Infrastructure as Code (Terraform) is a plus.
- Bachelor's or Master's degree in Computer Science, Information Technology, Engineering, or a related field.
- Databricks Certified Data Engineer Associate
- Databricks Certified Data Engineer Professional
- Microsoft Certified: Azure Data Engineer Associate (DP-203)
- Azure Fundamentals (AZ-900)
- Experience with real-time analytics.
- Knowledge of Lakehouse architecture.
- Experience with Agile/Scrum methodologies.
- Strong analytical and problem-solving skills.
- Excellent communication and stakeholder management abilities.
- Databricks
- PySpark
- Spark SQL
- Delta Lake
- Python
- SQL
- Azure/AWS/GCP (at least one cloud platform)
- ETL/ELT Development
- Data Lake Architecture
- Unity Catalog
- Delta Live Tables (DLT)
- MLflow
- Kafka/Event Hubs
- Azure Data Factory
- Terraform
- Azure DevOps/GitHub Actions
Similar Jobs
Fintech • Machine Learning • Payments • Software • Financial Services
Lead data engineering role responsible for designing, developing, testing, implementing, and supporting cloud-based data solutions. The position uses Python, Java, Scala, SQL, distributed computing, streaming technologies, data warehouses, NoSQL databases, and public cloud platforms. Responsibilities include collaborating with Agile and product teams, building scalable systems, conducting code reviews and unit tests, optimizing performance, tracking technology trends, and mentoring engineers.
Top Skills:
AgileSparkAWSCassandraDatabricksEmrGCPGurobiHadoopHiveJavaKafkaLinuxMapreduceAzureMongoDBMySQLNoSQLPythonRdbmsRedshiftScalaShell ScriptingSnowflakeSQLUnix
Healthtech • Information Technology • Software
Build, administer, and optimize enterprise Databricks data platforms in Azure. Responsibilities include platform operations, Terraform-based infrastructure automation, data ingestion enablement, governance through Unity Catalog, access controls, monitoring, troubleshooting, incident response, and reliability improvements. The role partners with data engineers, architects, analysts, infrastructure, and security teams to support scalable production analytics and data workloads.
Top Skills:
Azure Data LakeCi/CdDatabricksAzurePythonSnowflakeSQLTerraformUnity Catalog
Information Technology
Designs, operates, and scales AWS cloud infrastructure, networking, and Databricks platforms across multi-account environments. Manages Terraform-based infrastructure, CI/CD pipelines, EKS/Kubernetes clusters, enterprise data integrations, observability, incident response, and governance. Optimizes networking, compute performance, costs, reliability, and production data workloads while partnering with engineering teams on continuous delivery and platform operations.
Top Skills:
Amazon EksAmazon RedshiftAWSCi/CdCidrDatabricksDelta LakeGithub ActionsHelmIamIpv4JenkinsKubectlKubernetesMlflowSnowflakeTerraformUnity CatalogVpc
What you need to know about the San Francisco Tech Scene
San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.
Key Facts About San Francisco Tech
- Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
- Major Tech Employers: Google, Apple, Salesforce, Meta
- Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
- Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
- Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
- Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine


