ifm Logo

ifm

Data Engineer

Reposted Yesterday
In-Office
Sunnyvale, CA, USA
150K-450K Annually
Senior level
In-Office
Sunnyvale, CA, USA
150K-450K Annually
Senior level
Build and maintain large-scale data collection and preprocessing pipelines for NLP research. Develop web crawlers, APIs, and workflows; refine LLM outputs into structured datasets; collaborate with researchers to ensure data quality and document methodologies. Support scalable storage, retrieval, and distribution of datasets.
The summary above was generated by AI
About the Institute of Foundation Models
We are a dedicated research lab for building, understanding, using, and risk-managing foundation models. Our mandate is to advance research, nurture the next generation of AI builders, and drive transformative contributions to a knowledge-driven economy.

As part of our team, you’ll have the opportunity to work on the core of cutting-edge foundation model training, alongside world-class researchers, data scientists, and engineers, tackling the most fundamental and impactful challenges in AI development. You will participate in the development of groundbreaking AI solutions that have the potential to reshape entire industries. Strategic and innovative problem-solving skills will be instrumental in establishing MBZUAI as a global hub for high-performance computing in deep learning, driving impactful discoveries that inspire the next generation of AI pioneers.



The Role
 
As a Data Engineer specializing in Natural Language Processing (NLP) and large-scale data processing, you will quickly and effectively gather, curate, and prepare high-quality datasets to support cutting-edge NLP research. Your role will be instrumental in enabling researchers by delivering essential data through efficient and scalable engineering practices, including web crawling, LLM-generated content refinement, and robust data pipelines, primarily leveraging Python and related technologies.

Key Responsibilities

  • Rapidly collect, curate, and preprocess datasets based on detailed specifications provided by NLPresearchers,delivering data within tight timelines.
  • Develop and maintain efficient web crawling solutions, APIs, and automated workflows to continuously improve data collection processes.
  • Refine and evaluate outputs from Large Language Models (LLMs) to generate structured datasets suitable for model training and benchmarking.
  • Implement scalable data pipelines, ensuring efficient data processing, storage, retrieval, and distribution to research teams.
  • Collaborate closely with researchers and engineers to ensure collected data meets specified quality and relevance criteria.
  • Document data collection methodologies, dataset characteristics, and pipeline architecture clearly and effectively.
  • Engage with peer teams and participate in technical reviews to uphold best practices and data quality standards.
  • Represent MBZUAI at industry and research forums, showcasing technical capabilities in large-scale data processing and AI data infrastructure.

Academic Qualifications

  • Bachelor's degree in Computer Science, Data Science, Engineering, or a related technical field required
  • Master’s degree or PhD degree or equivalent experience in Computer Science, Data Engineering, or related technical fields preferred.

Professional Experience - Required

  • Extensive experience in data engineering, data processing, and automation using Python.
  • Demonstrated proficiency in designing and deploying web crawling solutions, automated data extraction, and processing pipelines.
  • Strong understanding of data structures, algorithms, databases, SQL, and performance optimization.
  • Experience working with cloud infrastructure and distributed data processing frameworks (e.g., AWS, Spark, Kafka, Kubernetes).
  • Excellent problem-solving abilities, attention to detail, and the capability to rapidly address technical challenges.
  • Strong communication and collaboration skills with cross-functional teams.

Professional Experience - Preferred

  • Proven track record of supporting NLP or AI research teams with rapid and reliable data delivery.
  • Experience working with large language models, including evaluation, efficient inference, and prompt engineering.
  • Experience with refining outputs from large-scale AI models, such as LLM-generated data.
  • Contributions to open-source projects, coding competitions, or high visibility in coding communities (e.g., GitHub, Stack Overflow).
  • Familiarity with the latest advancements in NLP data processing and large language model technologies.

Visa Sponsorship
This position is eligible for visa sponsorship.

Benefits Include
*Comprehensive medical, dental, and vision benefits 
 *Bonus
*401K Plan
*Generous paid time off, sick leave and holidays
*Paid Parental Leave
*Employee Assistance Program
*Life insurance and disability


Similar Jobs

An Hour Ago
Remote or Hybrid
United States
70K-120K Annually
Mid level
70K-120K Annually
Mid level
Cloud • Insurance • Payments • Software • Business Intelligence • App development • Big Data Analytics
Build, model, and maintain scalable BigQuery-based data solutions on GCP. Implement performant data models, storage partitioning/clustering, ETL improvements, data validation, and documentation. Collaborate with architects, data scientists, and engineers to deliver governed, high-quality data for reporting and AI initiatives.
Top Skills: Ansi SqlBigQueryBigquery SqlConfluenceGCPJIRAPythonSQL
Yesterday
Easy Apply
In-Office or Remote
6 Locations
Easy Apply
180K-230K Annually
Senior level
180K-230K Annually
Senior level
AdTech • Artificial Intelligence • Big Data • Machine Learning • Marketing Tech • Mobile • Software
Build and maintain scalable, reliable ML data platform systems to support dataset generation, model training, analytics, monitoring, and large-scale inference. Collaborate with ML, software, and infrastructure engineers to design cost-efficient data lake and training infrastructure using vendor and open-source tools to enable next-generation ML models.
Top Skills: GoPython
Yesterday
Easy Apply
Hybrid
San Francisco, CA, USA
Easy Apply
110K-130K Annually
Senior level
110K-130K Annually
Senior level
Fintech • Financial Services
Design, build, and maintain a high-quality data platform delivering private market data to internal and external clients. Implement features, maintain functionality, perform code reviews, unit tests, CI/CD deployments, and collaborate with Product, Design, and Engineering to deliver requirements on schedule.
Top Skills: AWSAzureC#Ci/CdData LakesGCPGitGitJavaJIRAKafkaLarge Language ModelsMedallion ArchitecturePostgresPythonSQLTypescriptWebsockets

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account