EXL Logo

EXL

Senior Data Engineer

Reposted 16 Days Ago
Remote or Hybrid
Hiring Remotely in United States
94K-154K Annually
Senior level
Remote or Hybrid
Hiring Remotely in United States
94K-154K Annually
Senior level
Design, build, and operate production-grade, event-driven data pipelines (Kafka/Flink) on GCP to deliver model-ready features. Optimize BigQuery SQL and Parquet performance, develop Python data workloads (Polars/Pandas), deploy ML pipeline components on Kubeflow/Vertex AI with Docker, design event store architectures, and collaborate with ML and platform teams while documenting architecture and standards.
The summary above was generated by AI

EXL is hiring a Senior Data Engineer to join a strategic AI / ML platform engagement with a leading specialty retailer. This is a hands-on build role embedded with the client's platform engineering team.

The role requires shipping production-grade data pipelines that feed real-time customer event data into machine learning workflows. The right person is comfortable owning the full lifecycle of pipeline design, build, and deployment: from streaming ingestion through event store design to model-ready feature delivery.

This is a high-visibility role with growth potential into a larger book of work as the engagement expands.

Salary Range: $93,900 - $154,200 annual base 

The posted range is the hiring range for this role — a subset of the broader range available to employees over time — and reflects base salary across our national hiring scale. Final offers are based on several factors, including the candidate's skills and experience, internal pay equity, work location, market conditions for the role, and the specific scope and responsibilities of the position. The top of the range is reserved for candidates who notably exceed the requirements; the lower end applies to those with less experience or fewer preferred qualifications. For positions based in higher-cost zones (e.g., California, New York, New Jersey), actual compensation may exceed the posted range; your recruiter will share specifics during the process.

Responsibilities
What You'll Do
  • Design and operate event-driven data pipelines using Kafka consumers and Flink jobs to process high-volume customer events (clicks, purchases, returns) in near-real time.
  • Build and optimize large-scale data transformations on Google Cloud Platform — BigQuery SQL, query performance tuning, and partitioning strategy at scale.
  • Develop Python data engineering workloads using Polars or Pandas at scale, with rigorous attention to Parquet partitioning, join performance on large datasets, and memory efficiency.
  • Build, deploy, and maintain ML pipeline components on Kubeflow Pipelines (KFP) and Vertex AI; package and deploy services with Docker.
  • Design event store architecture: partitioning by customer, time-ordered event assembly across heterogeneous sources, and schema management for mixed event types.
  • Partner with ML engineers, platform engineers, and data scientists to deliver clean, performant, model-ready data products.
  • Document architecture decisions and contribute to engineering standards across the platform team.
Qualifications
Required Skills & Experience
  • 6–12 years of experience in data engineering, platform engineering, or a closely related discipline.
  • Streaming: Production experience with Kafka consumers and Flink stream processing — building, deploying, and operating streaming jobs at meaningful scale.
  • GCP Data Stack: Strong SQL on BigQuery (or an equivalent cloud warehouse), with demonstrated query optimization, cost management, and partitioning chops.
  • Python Data Engineering: Hands-on with Polars or Pandas at scale; deep working knowledge of Parquet partitioning and performance on large joins.
  • ML Pipelines: Hands-on experience building and deploying components on Kubeflow Pipelines (KFP) and/or Vertex AI Pipelines; working proficiency with Docker.
  • Event Store Design: Demonstrated experience designing event stores — partitioning by customer, time-ordered event assembly across sources, schema strategy for mixed event types (clicks, purchases, returns).
  • Communication: Strong written and verbal communication; comfortable being the senior IC voice in design conversations with client stakeholders.
Nice to Have
  • Domain experience in Retail or E-commerce — customer journey data, transaction analytics, returns and exchanges modeling.
  • Exposure to schema registry tooling (e.g., Confluent), Iceberg, or Delta Lake.
  • Experience working in client-facing or consulting engagements.
  • Google Cloud certifications (Professional Data Engineer or equivalent).
Work Arrangement & Eligibility
  • This role requires 3–4 days per week onsite in Seattle, WA. Fully remote and out-of-state candidates will not be considered.
  • EXL is open to sponsoring H1B transfers for qualified candidates.

EXL Richmond, California, USA Office

Richmond, United States

Similar Jobs

Yesterday
In-Office or Remote
92K-164K Annually
Senior level
92K-164K Annually
Senior level
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Designs, develops, deploys, and supports scalable cloud data engineering solutions for HEDIS processing and healthcare quality reporting. Builds ETL pipelines, data integrations, SQL transformations, and AI-enabled workflow automation. Collaborates with stakeholders, auditors, and product teams to ensure data quality, regulatory compliance, and operational reliability. Troubleshoots production issues, supports cloud modernization, applies CI/CD and testing practices, and mentors engineers as a senior individual contributor.
Top Skills: Agentic AiAICi/CdDatabricksETLFhirGenaiAzureSQL
2 Days Ago
In-Office or Remote
7 Locations
168K-297K Annually
Senior level
168K-297K Annually
Senior level
Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency
Lead data modeling and ownership of critical pipelines supporting Block’s product data foundation. Build reliable, governed, and monitored datasets, data quality and lineage systems, experimentation infrastructure, and AI-assisted automation. Partner with product and engineering teams to translate business needs into end-to-end data solutions, participate in on-call support, and maintain pipeline SLAs.
Top Skills: AirflowDatabricksDbtGitOmniPrefectPythonSnowflakeSQLTerraform
3 Days Ago
Remote or Hybrid
7 Locations
168K-297K Annually
Senior level
168K-297K Annually
Senior level
Blockchain • Fintech • Mobile • Payments • Software • Financial Services
Build and optimize data models, pipelines, monitoring, lineage, and data quality systems supporting Block’s product data foundation. Own critical data engineering solutions across their full lifecycle, ensure pipeline reliability through on-call support, and translate business and product needs into automated data solutions. Develop AI-assisted workflows and agents for data quality, issue resolution, and operational efficiency while partnering with engineering, product, data science, and machine learning teams.
Top Skills: AirflowDatabricksDbtGitOmniPrefectPythonSnowflakeSQLTerraform

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account