Maximum of 25 job preferences reached.
Top Staff Data Engineer Jobs in San Francisco, CA
Artificial Intelligence • Fintech • Payments • Business Intelligence • Financial Services • Generative AI
Lead architecture, development, and operation of the company data and AI platform. Define technical strategy, build and mentor a high-performing data platform team, drive company-wide initiatives, deliver scalable multi-region data pipelines, oversee ETL frameworks, metrics stores, infrastructure and data security, and advocate best practices for lakehouse-style data processing and model governance.
Top Skills:
AirflowBigQueryDatabricksDataflowDataprocDbtGCPGoogle Cloud StorageKafkaKubernetesMcpRagSpark
AdTech • Marketing Tech
Lead architecture, build, and operate Prodege's high-scale data platform including batch, ELT, and near-real-time streaming pipelines. Deliver Medallion-modeled data, governance, lineage, observability, and feature/ML data infrastructure. Drive platform standards, optimize performance and cost, and mentor teams while partnering with product, analytics, and ML groups to enable analytics, experimentation, and AI-driven applications.
Top Skills:
DbtPythonSnowflakeSQL
Artificial Intelligence • Healthtech • Machine Learning • Biotech
Design and build petabyte-scale data pipelines and platform tooling to ingest, transform, catalog, and deliver heterogeneous biological datasets for AI research. Improve reliability, observability, and automation (agent-based curation/QA) while collaborating with AI researchers and scientists.
Top Skills:
SparkAWSAws CdkBamClaude CodeDatabricksFastqHdf5LlmsOme-ZarrOn-Prem InfrastructureRayTerraformVcf
Artificial Intelligence • Cloud • Information Technology • Machine Learning • Software • Sports • Analytics
Define and lead the technical architecture for a lakehouse on AWS, build production pipelines and data models for time-series and multi-source data, introduce data governance and catalog foundations, author platform playbooks (ADRs, runbooks, Terraform modules), deliver projects end-to-end, and mentor engineers while representing the platform to senior leadership and non-technical stakeholders.
Top Skills:
AWSDatabricksDelta LakeEmrGlueHudiIamIcebergLambdaPythonS3SnowflakeSparkTerraformTrino
Healthtech
Lead architecture and hands-on engineering to design, scale, and govern a cloud-native lakehouse using Databricks and Microsoft Fabric. Build batch/streaming pipelines, ensure HIPAA-compliant PHI handling, produce analytics-ready datasets, mentor engineers, and drive platform strategy, governance, and reliability.
Top Skills:
AzureCaboodleCi/CdClarityCogito CloudData FactoryDatabricksDbtDelta LakeEpicInfrastructure As CodeLakehouseMicrosoft FabricOnelakePysparkPythonSpark SqlUnity Catalog
Financial Services
Build and own the data platform backbone: design ingestion patterns, Snowflake warehouses, governance, and freshness strategy. Implement reliable batch and streaming pipelines, data quality, observability, and onboarding processes. Lead legacy migrations, create data contracts and SLAs, and apply AI-native tooling to automate debugging, documentation, and operational workflows while partnering with Analytics, Engineering, and vendors.
Top Skills:
AgentsAirbyteAirflowAzure Data FactoryDagsterDbtDbt CloudDebeziumElementaryFivetranIcebergLlmsMcp ServerMonte CarloPrefectPythonSnowflakeSQL
Financial Services
Lead design and ownership of a modern data platform: ingestion, Snowflake governance, pipeline reliability, migrations, data contracts, observability, and AI-enabled tooling. Partner with Analytics, Engineering, and vendors to deliver trustworthy data products and operate production pipelines.
Top Skills:
AgentsAirflowApache IcebergAPIsAzure Blob StorageAzure DevopsAzure Event HubsAzure SqlCdcCi/CdDagsterDbt CloudDbt TestsDebeziumElementaryLlmsMonte CarloSnowflake
Healthtech • Information Technology • Software • Telehealth
The role involves architecting and improving data systems, defining governance standards, optimizing performance, and mentoring engineers in data engineering.
Top Skills:
AirflowBigQueryDagsterDbtPythonRedshiftSnowflakeSQL
Automotive
The role involves building scalable data platforms, developing data pipelines, mentoring engineers, and enhancing data security while collaborating across teams.
Top Skills:
AirflowAWSAzureGCPKubernetesSparkTerraform
Machine Learning • Software
Own end-to-end client data migrations: reverse-engineer legacy ETL, design and build Airflow DAGs and Spark jobs, validate and reconcile data at scale, document architecture and decisions, and coordinate cross-functional teams through go-live and hypercare.
Top Skills:
Apache AirflowSparkAws S3ClaudeCursorGithub CopilotPythonSQLSQL ServerSsisT-Sql
Cloud • Digital Media • Enterprise Web • Marketing Tech • Software
Own and drive the architecture and roadmap of ClickUp's data platform. Build scalable, reliable data pipelines and AI/ML infrastructure using AWS serverless, Snowflake, dbt, and Terraform. Lead cross-team technical initiatives, optimize cloud costs, establish engineering standards (observability, testing, CI/CD), mentor senior engineers, and influence org-wide architecture decisions.
Top Skills:
AirflowAmazon KinesisAuroraAws FargateAws LambdaAws S3Aws Step FunctionsCdkCi/CdDagsterDbtDockerDynamoDBEmbedding PipelinesFeature StoreGitKafkaLlm FrameworksModel ServingPrefectPythonSnowflakeSQLTerraform
eCommerce • Fintech • Machine Learning • Retail
Lead architecture and execution of Faire's data platform and analytics capabilities. Build and operate scalable, secure ETL/ELT, streaming, storage, compute, orchestration, and data warehousing at petabyte scale. Drive performance, cost, monitoring, governance, and self-serve analytics; influence cross-functional teams and mentor engineers.
Top Skills:
AirflowAWSEtl/EltKotlinPythonSnowflakeSparkSQL
New
Track Smarter, Apply Better.
Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.
Use For Free
Consumer Web • Mobile
Design and implement scalable backend systems for Patreon's platform. Collaborate with engineers and product managers to maintain and optimize performance.
Top Skills:
MySQLNoSQLPythonRestful Apis
Automotive • Manufacturing
Lead product development and optimization of internal databases and MES for battery R&D and pilot manufacturing. Define product roadmaps, write software requirements and user stories, oversee third-party developers, design UI/data visualizations, architect scalable database structures, maintain data integrity, and serve as primary contact for training, troubleshooting, and documentation across engineering and manufacturing teams.
Top Skills:
ArbinBiologicFront-End Web FrameworksMesNewarePower BIPythonRelational DatabasesSQLTableau
Big Data • Machine Learning • Software • Analytics • Big Data Analytics
The role involves deep troubleshooting, root cause analysis, and architectural optimization in the Data and AI ecosystem to enhance platform reliability and supportability.
Top Skills:
Delta LakeHiveJavaPythonScalaSparkSQL
Fintech • Information Technology • Payments • Financial Services
Founding data engineer to design and build Payabli's data platform: architect lakehouse/warehouse, build batch and streaming pipelines, model canonical datasets, ensure data quality/observability, enforce access/masking/lineage for regulated financial data, and enable analytics/ML feature pipelines while establishing team standards and CI/CD.
Top Skills:
AirflowAWSAzureCdcDagsterDatabricksDbtDelta LakeEltFeature StoreFivetranFlinkGCPGreat ExpectationsKafkaKinesisMlflowMonte CarloOpenlineagePci DssPrefectPythonSnowflakeSparkSQLUnity Catalog
Information Technology • Sales
Build and scale the company's data infrastructure: design ingestion paths, event models, schema-flexible data models, change capture, query performance, freshness guarantees, and analytics serving for customer-facing dashboards and evaluation pipelines.
Top Skills:
BullmqChange Data Capture (Cdc)ClickhouseData WarehousesEvent PipelinesFlinkIcebergKafkaOlap SystemsPostgresRedisSparkTypesense
Artificial Intelligence • Machine Learning • Software • Cybersecurity
Design and implement a scalable data platform and integration pipeline to ingest, normalize, and index enterprise data. Build connectors, ETL/streaming pipelines, and storage/indexing strategies to support semantic search and GenAI use cases. Partner with applied AI teams, make build-vs-buy decisions, and create reusable patterns to scale integrations.
Top Skills:
AirflowApache FlinkApache KafkaData LakeDatabricksDockerKnowledge GraphsKubernetesLlmRetrieval Augmented Generation (Rag)Semantic SearchSQLTemporalTerraform
AdTech • Agency
Design, build, and operate large-scale data pipelines and ETL processes to normalize varied partner datasets. Lead technical decisions, mentor engineers, implement testing and monitoring, manage data warehouses/lakes, and collaborate cross-functionally to deliver robust data infrastructure that supports ML and product teams.
Top Skills:
SparkAWSData LakeData WarehouseDockerGithub ActionsTerraform
eCommerce • Hardware • Healthtech • Software
Design, build, and own end-to-end ingestion pipelines and dbt transformation layers on Databricks. Lead GCP->AWS/Databricks migration, validate parity, and decommission legacy systems. Implement CDC and batch patterns, ensure data quality and observability, document assets in Unity Catalog, and build reverse ETL integrations to operational systems. Partner with business stakeholders and write architectural decision records to define engineering patterns and acceptance criteria.
Top Skills:
AirflowAutoloaderAWSBigQueryCdcCloud FunctionsCloud RunDatabricksDbtDelta LakeDelta Live TablesGCPGreat ExpectationsLakeflowMulesoftNetSuitePysparkPythonSalesforceStripeUnity Catalog
Automotive
You will create large datasets and training recipes, develop methods for data mining, create scalable infrastructure solutions, and collaborate with ML infrastructure teams.
Top Skills:
C++Python
Automotive
Own and develop concurrent C++ backend services for Webviz to stream time-series and sensor data, integrate offboard storage and WebRTC, optimize latency and throughput, build APIs for automated triage/evaluation, plan technical roadmaps, and mentor engineers.
Top Skills:
BoqBorgC++CnsRpcSpannerWebrtc
Big Data • Machine Learning • Software • Analytics • Big Data Analytics
As a Staff Software Engineer, you will build distributed data systems, focusing on reliability and performance for large-scale data processing, using technologies like Apache Spark™ and Delta Lake.
Top Skills:
C++JavaScala
Fintech
Design, architect, and build cloud-native, large-scale data platforms and pipelines to support payments products, analytics, and AI/ML. Lead technical strategy, deliver production systems, collaborate cross-functionally, improve engineering standards, and support operations and incident resolution.
Top Skills:
AerospikeAppdynamicsAWSAzureCi/CdClouderaDatarobotElasticsearchGoogle Cloud PlatformHdfsIbm InfosphereKafkaKubernetesOraclePub/SubPycharmPysparkPythonRRabbitMQS3SagemakerScality S3ScikitSparkSQL ServerTensorFlow
Big Data • Information Technology • Software • Analytics
Design, build, and operate petabyte-scale data infrastructure: real-time ingestion, storage (open table formats), and serving. Architect scalable Iceberg-based data lakes, optimize Spark pipelines and streaming (Kafka/Flink), drive reliability, cost efficiency, and data contracts; collaborate with product and platform teams and establish best practices and tooling.
Top Skills:
AirflowAmazon S3Apache FlinkApache IcebergApache KafkaSparkAws GovcloudKubernetesPythonScala
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top San Francisco Companies Hiring Staff Data Engineers
See AllPopular Job Searches
All Filters
Total selected ()
No Results
No Results



.png)






.png)


.png)
















