Metriport Logo

Metriport

Senior Data Engineer

Posted 6 Days Ago
Hybrid
San Francisco, CA, USA
180K-220K Annually
Senior level
Hybrid
San Francisco, CA, USA
180K-220K Annually
Senior level
Own and scale the data platform: design and implement ETL/ELT pipelines, optimize warehouse/lake architecture, support streaming and batch ingestion, enable ML workflows, mentor engineers, write design documents, participate in sprints and on-call rotation, and deliver reliable clinical data at scale to customers.
The summary above was generated by AI
Senior Data Engineer

San Francisco, CA

Hybrid


About us

Metriport is an open-source data intelligence platform that helps healthcare organizations access and exchange patient data in real-time: record retrieval that used to take days now takes seconds, and clinicians walk into every appointment with the full patient picture. We integrate with all major US healthcare IT systems and tap into comprehensive medical data for 300+ million individuals.

We've found product-market fit with multi-million ARR, 100+ customers (including Amazon One Medical, Strive Health, Circle Medical, and Brightside Health), backing from top VCs, and years of runway. We're ready to scale. We're a tight-knit, high-performing team of mostly former founders (including two YC alumni). We're engineering-heavy, operate with minimal bureaucracy and high autonomy, and hire based on competence, not prestige. We push hard—founders work six days a week from our SF office—but give everyone freedom to craft their schedule. We measure output and we're committed to sustainable intensity.

About you

We're looking for a data engineer who can operate at a Senior level:

  • You've built and scaled data pipelines and systems, with hands-on experience across the ecosystem — distributed processing, lakehouses, warehouses, streaming, orchestration. You know when each is (and isn't) the right tool for the job, and people usually come to you for guidance on how to move, transform, and serve data reliably.

  • You fully own your work — technical decisions, delivery, results — and you level up the engineers around you through code reviews, pairing, and guidance. Multiplying the team's output energizes you as much as shipping your own.

  • You're entrepreneurial-minded with an olympian-level work ethic (about half our engineering team are former founders).

  • You understand data reliability, quality, and governance as first-class parts of delivery.

  • You care about delivering value to customers, not about what frilly new tech is under the hood.

  • When someone scopes a project for 3 weeks, you ask "why can't it be done in 3 days?" — and you help others develop that same instinct.

  • You're a hacker at heart, with a good sense of which rules should, and shouldn't, be broken.

What you'll be doing

We ingest clinical data for millions of patients from external healthcare sources, with continuous updates for a growing subset of those patients. You'll be a catalyst to scale the data platform that powers our product — and ship it to customers fast.

Day to day, that looks like:

  • Raising the technical bar for our data platform: building on and improving our warehouse, data lake, and ETL/ELT architecture so it scales with patient and customer growth, and helping evaluate the right tools (batch and streaming processing, table formats, orchestration, query engines).

  • Driving data projects end-to-end: writing Design Documents, shipping v0's quickly, and iterating to v1 and beyond.

  • Supporting AI/ML efforts - making sure the AI Engineers have the data they need.

  • Multiplying the team: mentoring engineers on data fundamentals, reviewing designs and PRs, and judging when to invest in quality vs. ship fast.

  • Participating in bi-weekly sprint planning and retros, joining our daily 30-min remote stand-up at 7:30am PST (our only mandatory meeting), and taking part in the on-call rotation.

Example projects you could own:

  • Scaling our patient data consolidation pipeline (deduplication, normalization, hydration) to handle 100x today's volume without 100x the cost.

  • Building pipelines that deliver clinical data directly into customers' data warehouses, reliably and at scale.

  • Building the ingestion path for customers pushing large volumes of their own data into the platform.

  • Building document-processing pipelines that extract structured data from PDFs, images, and free text to feed ML models.

Requirements
  • 6+ years of engineering experience, with a heavy lean towards data engineering — building, maintaining, and scaling pipelines processing terabytes of data and millions of events a day.

  • Experience across the data stack — ingestion, storage, processing, warehousing, serving — and an understanding of the tradeoffs (cost, latency, correctness, operability) at each layer.

  • Experience with modern, cloud-native data stacks: e.g., Spark, open table formats (Parquet, Iceberg, Delta) on S3, and warehouses (Snowflake, BigQuery, Redshift).

  • Strong software engineering fundamentals — you write production code (we're a TypeScript shop, with Python in data/ML workflows), not just orchestration configs.

  • Experience mentoring or guiding other engineers — through code reviews, pairing, design feedback, or onboarding.

  • Located in San Francisco / Bay Area, or willing to relocate.

  • Bonus:

    • Experience with streaming systems (Kafka, Kinesis), dbt, or orchestration tooling (Airflow, Dagster).

    • Experience building or supporting ML/data science workflows (feature pipelines, model inputs/outputs, unstructured data extraction).

    • Healthcare standards/technologies: FHIR, HIE, IHE, EHR/EMR, NPI, TEFCA, ADT, HL7, HEDIS, RAF, SNOMED, LOINC, ICD-10, etc.

Benefits
  • Competitive equity + compensation package 🚀

  • Full family Platinum health insurance, dental, and vision coverage 🦷

  • 401(k) retirement plan + matching 💰

  • Flexible work from home or in-office 🏢

  • Healthy lunches are complimentary when working in-office (and breakfast + dinners as needed) 🍏

  • Quarterly company off-sites with the team ⛷️

  • MacBook provided by us 💻

  • Unlimited PTO (we work hard, but trust you to take time you need to be at your best) 🧘‍♂️

Our tech

Core business logic in Node.js and TypeScript, with Python in data and ML workflows. AWS across the board (ECS, Lambda, SQS, SNS, Batch, etc.), infrastructure as code with CDK. Data lives in S3, PostgreSQL/Aurora, DynamoDB, Snowflake, and our FHIR server — with Athena for querying S3 and SageMaker for ML. Our data platform is still early: you'll shape what we adopt next, picking the best tool for the job rather than the trendiest one.

Metriport provides equal employment opportunities (EEO) to all employees and applicants for employment without regard to race, color, religion, sex, national origin, age, disability, genetics, sexual orientation, gender identity, or gender expression. We are committed to a diverse and inclusive workforce and welcome people from all backgrounds, experiences, perspectives, and abilities.

HQ

Metriport San Francisco, California, USA Office

San Francisco, CA, United States

Similar Jobs

9 Days Ago
Hybrid
124K-280K Annually
Senior level
124K-280K Annually
Senior level
Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI
Lead development and delivery of data engineering solutions for clients: design data architectures, build and maintain ETL pipelines, apply advanced analytics, implement BI visualizations, coach teams, manage managed-services SLAs, and engage stakeholders to validate and refine outcomes.
Top Skills: AWSDatastageDb2ETLJavaOracle BiPythonQlikviewRedshiftSQL Server
9 Days Ago
Hybrid
99K-232K Annually
Senior level
99K-232K Annually
Senior level
Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI
Lead data engineering engagements delivering ETL/ELT solutions, pipelines, and cloud analytics. Build and maintain architectures (DataStage, AWS/Redshift, DB2/SQL Server, GoldenGate/CDC, S3/Glue), oversee BI and visualization, manage client accounts and teams, mentor junior staff, ensure project delivery and quality, and drive process and technology adoption.
Top Skills: AWSBirtCdcDatastageDb2Etl/EltGlueGoldengateJavaPerformance TuningPipeline ArchitecturePythonQlikviewRedshiftS3SpotfireSQL Server
Yesterday
Hybrid
5 Locations
72K-212K Annually
Senior level
72K-212K Annually
Senior level
Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI
Design and implement AI systems and scalable machine learning models, perform data wrangling and software engineering to deploy AI solutions, interpret data for client recommendations, collaborate across teams, uphold technical standards, and build client relationships while navigating complex, ambiguous problems.
Top Skills: AIAWSDatabricksGCPMachine LearningAzureSnowflake

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account