Baselayer Logo

Baselayer

Data Engineer

Posted 3 Days Ago
In-Office
San Francisco, CA, USA
120K-150K Annually
Junior
In-Office
San Francisco, CA, USA
120K-150K Annually
Junior
Build and maintain production ETL/ELT pipelines ingesting public records, web signals, and fraud telemetry. Develop data models and transformation layers using Dataflow, Spark, and Airflow; implement data quality, observability, and alerting; optimize warehouse performance, freshness, and cost; and collaborate with data science, ML, and product teams. Support security and regulatory standards for sensitive data, document systems, and communicate with technical and non-technical stakeholders.
The summary above was generated by AI

ABOUT BASELAYER

Every business in America needs a bank account to exist. The system that decides whether they're real, who's behind them, and whether they're a risk, runs on infrastructure from the 1980s. We're rebuilding that layer from scratch.

Baselayer is the identity layer for institutions across the United States — the most complete business graph in America and every human tied to it. We fuse public records, IRS data, sanctions lists, web signals, and fraud telemetry from 2,200+ financial institutions into a single graph that resolves any business and the humans behind it in milliseconds. The legacy credit bureaus took 50 years to build something that gets 60% match rates. We've built something that gets 98% in under two years.

Today we're trusted by over 20% of financial institutions in America — including FIS, Rho, Socure and leading loan infrastructure providers. But the graph is becoming infrastructure for anyone who needs to know if a business is real and worth trusting: gig platforms, marketplaces, AI companies, and commerce infrastructure at scale.

Trust is the substrate of every financial transaction. We're rebuilding it.

ABOUT THE TEAM

We're solving real-time entity resolution at a scale no one else has cracked — fusing dozens of data sources into a single business identity graph and resolving any entity in milliseconds. It's a graph AI problem, a retrieval problem, and a fraud-modeling problem stacked on top of each other. The technical depth is real.

You'd be joining a small team where the data moat is defensible, the research problems are open, and the infrastructure you build becomes load-bearing for businesses. Ownership is real. Velocity is real. There's no layer of process between an idea and shipping it.

We're at an inflection point — the graph is built, the match rates speak for themselves, and the hardest problems are still ahead: graph embeddings, fraud propagation models across the business network, real-time traversal at sub-100ms latency, and expanding the identity layer beyond finance into every platform that needs to trust a business.

If you want to work on something foundational — the kind of infrastructure that gets built once and everything else runs on top of — this is it.

ABOUT THE ROLE

Baselayer is building the most comprehensive, accurate, and continuously-current identity graph of US businesses — fusing public records, IRS data, sanctions lists, web signals, and fraud telemetry from thousands of financial institutions into a single graph that resolves any business in milliseconds. None of that works without world-class data infrastructure. We’re hiring a Data Engineer to help build and run the pipelines and models that turn messy, heterogeneous data into trustworthy, production-grade signal. You’ll write real production code in your first weeks, own pipelines end to end, and learn alongside senior data and ML engineers who will invest in your growth. This is a role for an early-career engineer who wants to be close to the action: feeding the models, not just cleaning up after them.

WHAT YOU'LL DO

  • Build and maintain ETL/ELT pipelines that ingest and normalize public records, web signals, and fraud telemetry from dozens of sources
  • Develop data models and transformation layers (Dataflow, Spark, Airflow) that power fraud detection, KYB, and customer-facing APIs
  • Implement data quality checks, observability tooling, and alerting so problems surface before customers see them
  • Tune pipelines and queries for performance, freshness, and cost in our cloud data warehouse
  • Work with data scientists, ML engineers, and product to make clean, well-modeled data available for entity resolution and scoring
  • Help ensure pipelines meet security and regulatory standards for sensitive data (SOC 2, GDPR, KYC/KYB)
  • Document what you build and translate between technical and non-technical stakeholders so the rest of the team moves faster

MINIMUM REQUIREMENTS

  • 1+  years of experience in data engineering, working with Python, SQL, and cloud-native data platforms
  • Experience building and maintaining ETL/ELT pipelines in a production environment
  • Working knowledge of modern data stack tooling (e.g. Dataflow, Spark, Airflow or equivalents)
  • Hands-on experience with cloud data warehouses or lakes (e.g. BigQuery, Snowflake, or equivalents)
  • Solid data modeling fundamentals and real care for data integrity and reliability
  • Comfort with both structured and unstructured data, and a feel for what clean, scalable architecture looks like

WHAT SETS YOU APART

  • Curiosity about AI/ML infrastructure and a desire to be close to the models, not just the cleanup after them
  • Experience with streaming or real-time data systems (e.g. Kafka, Pub/Sub)
  • Exposure to KYC/KYB, fraud, risk, or underwriting data, and the ethical care that sensitive information demands
  • GCP experience (BigQuery, Cloud Run, Dataflow, Pub/Sub)
  • You care deeply about data quality and trust, and build systems others can rely on
  • You’ve worked without a playbook before, and you take direct feedback well and act on it fast

WORK LOCATION

  • Based in SF; hybrid - 4 days per week in office.

COMPENSATION

  • Salary Range: $120,000 – $150,000 + Equity

BENEFITS

  • Time off when you need it: Flexible PTO so you can recharge without red tape.
  • In-person energy: We're based in SF and meet in the office 4 days a week.
  • Competitive compensation: We pay well and back it with equity. We want you to think and act like an owner.
  • Career rocket fuel: You'll help build the foundation of a high-growth startup, working side by side with experienced founders and team members who've done it before.
  • Benefits on us: We cover 100% of your health, dental, and vision premiums. No surprise deductions from your paycheck.
  • 401(k) with company match: We match your contributions so your future self benefits too
  • HSA contributions included: We contribute to your HSA on applicable plans, so your coverage works as hard as you do
  • Stay healthy, stay sharp: A $250 monthly gym stipend to help you bring your best self to work, and everywhere else
  • A seat at the table: We believe in transparency, radical candor, and giving every team member a voice 🔥

Similar Jobs

2 Days Ago
Hybrid
100K-130K Annually
Mid level
100K-130K Annually
Mid level
AdTech • Consumer Web • Digital Media • eCommerce • Marketing Tech • SEO
Designs and maintains data collection infrastructure, resilient pipelines, ETL/ELT processes, data warehouses, streaming architecture, BI reporting, and analytical systems. Troubleshoots technical defects, performs root-cause analysis, maintains specialized datasets, and owns end-to-end data collection, extraction, and cleansing. The role requires experience with AWS, Google Cloud, ETL tools, programming languages, NoSQL databases, and BI platforms.
Top Skills: Amazon DocumentdbAmazon KinesisAmazon RedshiftAmazon S3Amazon SqsAWSBigQueryCloud Data FusionDruidDynamoDBGoogle AnalyticsGoogle Cloud PlatformGoogle Cloud StorageHiveInformaticaJavaMongoDBPower BIPrestoPythonSsisTableauTalend
2 Days Ago
Remote or Hybrid
United States
160K-260K Annually
Expert/Leader
160K-260K Annually
Expert/Leader
Artificial Intelligence • Cloud • Payments • Software • Business Intelligence • Generative AI • Automation
Define and govern enterprise-scale data architecture across batch, streaming, warehouse, lakehouse, transactional, and AI use cases. Establish standards for data quality, lineage, access, cataloging, governance, observability, and SLAs. Architect AI-enabled workflows, resolve complex architecture issues, influence roadmaps, and mentor engineers through hands-on technical leadership. The role requires 15+ years of software, data engineering, or architecture experience and expertise in large-scale data platforms and modeling.
Top Skills: AIBatch ProcessingBigQueryData CatalogsData WarehousesDbtFeature StoresGCPLakehousesOlapOltpStreaming ArchitecturesVector Stores
5 Days Ago
In-Office
San Francisco, CA, USA
140K-210K Annually
Senior level
140K-210K Annually
Senior level
Fintech • Information Technology • Payments • Sharing Economy • Financial Services • Cryptocurrency
Designs, builds, administers, and supports cloud data architectures, databases, schemas, metadata repositories, ETL processes, and analytic models. Leads data quality, governance, troubleshooting, testing, automation, reporting, dashboards, and infrastructure design. Collaborates with Agile teams to translate business requirements into technical capabilities and supports decision-making for Audit and business management. The role is full-time onsite at an eligible Federal Reserve location and requires extensive security screening.
Top Skills: Amazon Aurora PostgresqlAmazon EbsAmazon S3AribaAWSAws GlueAws IamAws LambdaAws MonitoringAws NetworkingGitlab Ci/CdJavaJava EePythonScalaSQLTerraformWorkday AdaptiveWorkday Prism

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account