Maximum of 25 job preferences reached.
Top Staff Data Engineer Jobs in San Francisco, CA
Transportation
Build scalable ML data pipelines for Waabi's autonomous driving platform. Design, optimize, and manage datasets and training processes while collaborating with scientists and engineers.
Top Skills:
Apache AirflowApache BeamApache HadoopSparkAws Step FunctionsGoogle Cloud DataflowJaxPythonPyTorchTensorFlow
Cloud
Design and implement data-intensive platform components, mentor engineers, and optimize streaming infrastructure to enable scalable data services at Okta.
Top Skills:
AWSBeamFlinkHadoopJavaKafkaKinesisKubernetesSnowflakeSpark
Artificial Intelligence • Enterprise Web • Healthtech • Software
Build and operate the end-to-end data layer: ingest and sync EHRs and customer systems, transform messy sources into a normalized patient/provider model, maintain data freshness and query performance, and enable natural-language query interfaces and serving for production AI workloads.
Top Skills:
Apache IcebergAPIsCdcDatabricksDelta LakeEhrsEtl/EltEvent-Driven ArchitectureFhirHl7KafkaNatural Language QueryReverse EtlSnowflakeSparkSQLUnity Catalog
AdTech • Marketing Tech
Lead design and implementation of identity resolution and data governance systems. Build batch and streaming pipelines, APIs, and services to ingest, normalize, link, and version identity data. Ensure auditable matching, data lineage, quality checks, schema enforcement, access controls, and privacy-by-design. Collaborate cross-functionally with product, analytics, legal, privacy, and security teams and create documentation, monitoring, and runbooks.
Top Skills:
SparkAPIsAWSBatch And Streaming Data PipelinesCloud Data WarehouseData LakeScalaStorage Formats
Social Media
Lead design and operation of identity resolution and data governance systems. Build batch and streaming pipelines, APIs, and tooling for identity ingestion, matching, lineage, quality, access controls, and privacy compliance (GDPR/CCPA). Collaborate with product, analytics, legal, and security to ensure privacy-by-design, monitoring, and documented runbooks.
Top Skills:
APIsAWSBatchCloud WarehousesData LakesScalaSparkStreaming
Social Media
Lead architecture and strategic direction for Pinterest's data warehouse, analytics tools, and data governance. Drive cross-functional initiatives to build scalable data platforms, AI-assisted pipeline tooling, and analytics capabilities. Mentor engineers, define policies and tooling, and deliver measurable adoption and business impact across the company.
Top Skills:
AirflowClaude CodeCodexCursorDatahubExadataFlinkQuerybookSparkSupersetTrino
Fintech • Real Estate
Lead design and delivery of customer-facing data product applications, building full-stack experiences (frontend, backend, APIs) using TypeScript, React, Node.js, and Python. Partner with Product and data teams to convert analytics and prototypes into production-grade workflows, establish software engineering best practices, mentor engineers, and prioritize customer value through testing, observability, and operational ownership.
Top Skills:
Node.jsPythonReactTypescript
Information Technology • Software • Automation
As a Staff Software Engineer, you will design high-performance data infrastructure, develop scalable backend systems, and collaborate on advanced engineering solutions.
Top Skills:
Apache ArrowApache FlinkArgocdAWSDruidGoKafkaKubernetesParquetPinotPostgresPrometheusPythonRustTerraformTimescaledb
Blockchain • Software
As a Staff Software Engineer, you'll lead the architecture and implementation of core systems and data platforms, ensuring adaptability and planning technical strategies.
Top Skills:
AIEmrMapreduceMlMongo ShardingPrestoSparkTrino
Healthtech
As a Staff Engineer, you will lead backend and data platform systems for Trial Library, ensuring reliability and scalability while utilizing AI tools to enhance operational efficiency in clinical trial navigation.
Top Skills:
AWSBedrockDrizzleFargateLambdaPostgresPythonRdsReactSqsTerraformTypescript
Marketing Tech
Lead technical design and hands-on development of Pantheon’s data platform, extending ingestion, curation, retention services and ML/LLM infrastructure. Mentor engineers, drive architecture and engineering standards, operate in a DevOps model, partner cross-functionally, and participate in on-call support to ensure reliable, compliant data access and analytics.
Top Skills:
AirflowBigQueryCi/CdDockerEmbedding/Vector PipelinesFeature StoreFireboltGoogle Cloud PlatformGreat ExpectationsKubernetesPythonRedshiftSnowflakeSnowpark For PythonTerraformYaml
Blockchain • Financial Services • Cryptocurrency • Web3 • Quantitative Trading
Lead and implement architecture for Ripple's Caspian Data Platform, owning ingestion, transformation, governance, and data quality. Design Databricks pipeline patterns, build critical components, drive AI tooling integration, mentor engineers, and lead complex cross-team initiatives from problem statement to production delivery.
Top Skills:
AWSDatabricksDelta LakeDelta Live TablesLlmsPythonSparkSQLUnity Catalog
New
Track Smarter, Apply Better.
Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.
Use For Free
Software
Build and evolve Ridgeline's AI data platform: design secure API integrations and MCPs, implement RAG pipelines and vector database solutions, scale distributed systems, enforce data lineage/access controls, and mentor peers while shaping AI platform strategy.
Top Skills:
Api GatewaysAWSClaude CodeCursorModel Context Protocols (Mcps)Oauth 2.0PgvectorPineconePythonRag PipelinesSnowflakeSsoTerraformVector DatabasesWeaviate
Artificial Intelligence • Information Technology • Machine Learning
Build and scale AI data and ML infrastructure: evaluation systems, fine-tuning pipelines, agent-first product surfaces, high-throughput data workflows, and integrations of new models into production.
Top Skills:
GCPGraphQLJavaKafkaKotlinKubernetesMySQLNode.jsPostgresPubsubPythonReactReduxSpannerTypescript
Artificial Intelligence • Software
The Staff Engineer will lead technical direction on data infrastructure, focusing on ingestion frameworks, storage architecture, and orchestration to support scientific workflows. Responsibilities include mentoring, ensuring reliability, and collaborating with cross-functional teams to establish coding standards and data integrity.
Top Skills:
AirflowAWSDagsterDelta LakeFlyteHudiIcebergKubernetesPythonSQL
Automotive
Lead development of data pipelines powering ML for motion planning and decision making. Prototype pipeline and model improvements, fine-tune and evaluate models, deploy and validate models on fleet with monitoring, collaborate with perception/research/simulation teams, and mentor engineers in data-driven reasoning approaches.
Top Skills:
Api DesignData PipelinesFleet Deployment And MonitoringLarge-Scale Models And DatasetsMachine LearningModel Evaluation WorkflowsModel Fine-Tuning
Big Data • Machine Learning • Software • Analytics • Big Data Analytics
Develop post-training recipes and systems for enterprise data agents capable of autonomous planning, code generation, and multi-step workflows. Partner with product teams to turn prototypes into production features for Databricks' Genie, and build context-discovery systems that use lakehouse data, notebooks, and code. Provide design, execution, debugging, and mentorship to raise team technical standards.
Top Skills:
Agentic RlAgentsGenieLakehouseLlmsModel TrainingNotebooksPost-Training WorkflowsReinforcement LearningRl Environments
Fintech • Analytics • Financial Services
Lead architecture and scaling of the Databricks data warehouse and self-service data platform. Build reusable batch and streaming frameworks, enforce governance, SLOs, lineage, observability, CI/CD, and multi-tenancy. Mentor engineers, run design reviews, and ship documented patterns to enable analytics, data science, and app teams to move quickly without blocking on platform tickets.
Top Skills:
AugmentAWSAzureAzure Event HubCi/CdClaudeDatabricksDbtGCPInfrastructure-As-CodePythonTerraform
Mobile • Software
Design, build, and lead a scalable, configuration-driven data integration platform. Improve performance, automate monitoring and deployments, provide escalated support, and collaborate across product and engineering to deliver reliable, secure data integrations.
Top Skills:
AWSAws Api GatewayAws LambdaCoalescePythonSnowflakeSQL
Aerospace
Design, implement, and maintain scalable distributed systems and storage for petabyte-scale 3D geospatial data. Build high-performance APIs, data ingestion and processing pipelines, monitoring/operational tooling, and collaborate with cross-functional teams to define requirements and drive the backend roadmap.
Top Skills:
AWSAzureC++DockerGCPJavaKubernetesPythonTerraform
Information Technology • Software
Lead architecture and technical direction for data integrations moving clinical and financial data at scale. Design orchestration, pipelines, data stores, observability, security/compliance, and mentor engineers while owning end-to-end cross-system initiatives and reliability.
Top Skills:
AthenahealthCernerContainer RuntimeDistributed TracingDnsEpicFhirGCPGoHipaaHl7 V2KnativeKubernetesLoggingMonitoringNatPostgresPythonRedisSoc 2Static IpsTemporalTls
Artificial Intelligence • Software • PropTech • Generative AI
Lead architecture and implementation of agentic AI systems powering Zuma's platform. Rebuild onboarding and integration frameworks, implement analytics and observability for LLM performance, translate product needs into technical solutions, mentor engineers, establish engineering standards, and work directly with customers to enable scalable, production-grade AI experiences.
Top Skills:
Agentic FrameworksFine-TuningLlmsNode.jsPythonReact
Edtech
Lead design and delivery of AI-powered, data-driven workflows and foundations: build retrieval/evidence pipelines, durable job execution, monitoring/evaluation, data contracts, identity resolution, and reusable abstractions while partnering with product and engineering to measure adoption and impact.
Top Skills:
AuditabilityData ContractsData WarehouseDeduplicationEmbedding PipelinesEvent-Driven SystemsIdentity ResolutionJob QueuesLakehouseLineage/Access ControlsLlmsLoggingObservability/MonitoringPii HandlingRelational DatabasesRetrieval PatternsSchedulersTool CallingVector Indexing
Artificial Intelligence • Information Technology • Machine Learning • Software
The role involves designing and scaling the data platform for Actively AI, including integrating data from various sources, ensuring data quality, and supporting real-time decision-making for automated agents.
Top Skills:
AirflowBigQueryDataflowDbtFivetranKafkaPub/SubPythonSnowflakeSQL
Automotive • Robotics • Software • Transportation
Build and automate large-scale ML data pipelines and tooling for autonomous trucking. Improve dataset curation, deduplication, labeling, evaluation, and metrics infrastructure. Collaborate with robotics, autonomy, and infra teams to support continuous learning, deployment, and scalable AI infrastructure for Kodiak's fleet.
Top Skills:
Apache AirflowData LakeEltLlmsMetaflowPythonSQL
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top San Francisco Companies Hiring Staff Data Engineers
See AllPopular Job Searches
All Filters
Total selected ()
No Results
No Results













.png)


















