Procore Technologies Logo

Procore Technologies

Staff ML Data Engineer (Datagrid)

Reposted 9 Days Ago
In-Office
San Francisco, CA, USA
227K-313K Annually
Senior level
In-Office
San Francisco, CA, USA
227K-313K Annually
Senior level
Lead design and build of scalable batch and streaming data pipelines for multimodal and spatial AI. Enable dataset curation, versioning, lineage, and data quality for research-to-production ML workflows. Partner with researchers and architects, run proofs‑of‑concept, optimize cost and performance, and mentor engineers to maintain observable, reproducible AI data systems.
The summary above was generated by AI

We’re looking for a Staff ML Data Engineer to join Procore’s AI & Frontier Models organization. In this role, you’ll be responsible for designing and building the data systems that power frontier‑scale machine learning research and applied AI products, with a particular focus on spatial intelligence and multimodal data. The primary goal of this role is to ensure that researchers and engineers can reliably discover, curate, transform, and operate on large‑scale datasets that move from experimentation to production.

As a Staff ML Data Engineer, you’ll work closely with ML researchers, applied ML engineers, and system architects to turn ambiguous research needs into scalable, production‑ready data pipelines. You’ll remain deeply hands‑on while providing technical leadership in data architecture, quality, and operational excellence. This is an opportunity to shape how Procore builds, evaluates, and deploys frontier models by ensuring the underlying data systems are robust, observable, and designed for iteration.

This role reports reports into the Manager, Software Engineering, and is based in our San Francisco office, supporting Procore's Datagrid AI Division. Given the collaborative and fast moving nature of this work, we are seeking candidates who are available to work onsite in a hybrid model at a minimum of 3 days per week. This is an immediate opening!

What you’ll do
  • Act as the technical lead for data engineering efforts supporting frontier model research and applied ML systems.

  • Design, build, and maintain scalable batch and streaming pipelines for multimodal data (e.g., documents, images, spatial metadata).

  • Partner closely with researchers and architects to translate experimental workflows into reliable, repeatable data systems.

  • Lead the development of dataset curation, versioning, and lineage workflows that support rapid experimentation and reproducibility.

  • Establish and uphold standards for data quality, validation, observability, and cost efficiency across AI data pipelines.

  • Contribute to data architecture decisions spanning research environments and production systems.

  • Identify gaps or inefficiencies in existing data workflows and run proofs‑of‑concept to evaluate improvements.

  • Mentor other engineers through code reviews, design discussions, and hands‑on collaboration.

What we’re looking for
  • Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience.

  • 8+ years of experience designing and operating complex data systems in production or research‑adjacent environments.

  • Strong proficiency in SQL and Python; experience with data‑intensive or distributed systems.

  • Proven experience building scalable data pipelines that support machine learning training, evaluation, or inference workflows.

  • Solid understanding of data modeling, dataset lifecycle management, and data quality best practices.

  • Comfort operating in highly ambiguous problem spaces and collaborating closely with researchers and architects.

  • Demonstrated ability to lead through direct technical contribution, mentorship, and setting engineering standards.

  • Strong communication skills, with the ability to explain technical tradeoffs to both research and engineering audiences.

Nice to have experience with technologies such as:

  • ML & Research Data: Large‑scale dataset curation, annotation workflows, experiment tracking, reproducibility tooling

  • Data Platforms: Databricks, Spark, lakehouse architectures, cloud data warehouses

  • Streaming & Pipelines: Kafka, Pub/Sub, event‑driven data architectures

  • Orchestration & Observability: Airflow, Dagster, data quality and lineage tools

  • Cloud & Infrastructure: AWS or GCP, containerized data workloads, CI/CD, infrastructure‑as‑code

  • Performance & Cost: Optimizing data pipelines for GPU‑backed training and large‑scale inference workloads

Additional Information

Base Pay Range:

227,332.00 - 312,581.50 USD Annual

This role may also be eligible for Equity Compensation and/or Bonus Incentive Compensation. Procore is committed to offering competitive, fair, and commensurate compensation. Actual compensation will be based on a candidate’s job-related skills, experience, education or training, and location.

For Los Angeles County (unincorporated) Candidates:

Procore will consider for employment all qualified applicants, including those with arrest or conviction records, in accordance with the requirements of applicable federal, state, and local laws, including the City of Los Angeles’ Fair Chance Initiative for Hiring Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act.

A criminal history may have a direct, adverse, and negative relationship on the following job duties, potentially resulting in the withdrawal of the conditional offer of employment: 1. appropriately managing, accessing, and handling confidential information including proprietary and trade secret information, as well as accessing Procore's information technology systems and platforms; 2. interacting with and occasionally having unsupervised contact with internal/external customers, stakeholders, and/or colleagues; and 3. exercising sound judgment.

Similar Jobs

17 Minutes Ago
Remote or Hybrid
United States
42K-57K Annually
Junior
42K-57K Annually
Junior
Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Administers group insurance transactions for participants, plan sponsors, customers, and internal partners. Reviews and enters enrollment data, manages eligibility and beneficiary records, handles plan referrals, maintains automated systems, troubleshoots data issues, develops customer relationships, and researches escalated cases. The role also supports process improvements, explains procedures during customer interactions, and ensures accurate, timely service delivery.
Top Skills: Microsoft Office Suite
45 Minutes Ago
Easy Apply
Hybrid
San Francisco, CA, USA
Easy Apply
105K-105K Annually
Senior level
105K-105K Annually
Senior level
Artificial Intelligence • Cloud • Security • Software • Cybersecurity
The Senior Mid-Market Account Executive will engage with mid-to-large size customers to drive new business growth through strategic sales efforts and collaboration.
Top Skills: DiscoverorgSales Navigator
47 Minutes Ago
Easy Apply
Hybrid
San Francisco, CA, USA
Easy Apply
66K-88K Annually
Junior
66K-88K Annually
Junior
Artificial Intelligence • Cloud • Security • Software • Cybersecurity
Premier Support Engineers at Datadog engage with clients to provide technical support, develop relationships, and drive product discussions while working in a fast-paced environment.
Top Skills: Datadog IntegrationsLinuxSaaS

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account