Mind Robotics Logo

Mind Robotics

Data Infrastructure Engineer

Reposted 4 Days Ago
In-Office
Palo Alto, CA, USA
Entry level
In-Office
Palo Alto, CA, USA
Entry level
The role involves building large-scale data systems for robotics AI training, focusing on data pipelines, storage, and workflows.
The summary above was generated by AI
About Mind:

Mind Robotics is building Physical AI for real-world industrial deployment, starting with the factory floor. We believe the hardest problems in AI are solved when researchers and engineers are hands-on with the physical world every day - and we're looking for people who are passionate about robotics, value ownership, and are excited to tackle difficult problems. Join us if you want to move beyond digital intelligence and put intelligence into motion.

About the team and the role:

Mind Robotics is building robots that learn from real-world experience. That starts with a data engine: the pipelines and infrastructure that turn raw, messy, multimodal sensor streams from robots and human demonstrations into high-quality, well-curated training data at scale.

As a Data Infrastructure Engineer, you'll build and operate the pipelines that turn raw sensor and demonstration data into training-ready datasets — from ingestion off real robots and capture devices, through processing and quality filtering, to the dataloaders that feed model training. The systems work today; your job is to build within the architecture, take ownership of specific pipelines and services, and help harden the system as it scales.

Responsibilities:
  • Build and scale ingestion pipelines for high-volume, high-dimensional sensor data.

  • Build and maintain data quality and curation systems, including both manual review workflows and auto-labeling.

  • Contribute to storage and retrieval design for large-scale multimodal datasets, including format choices and versioning/lineage.

  • Build and operate distributed processing infrastructure (batch and streaming).

  • Debug production issues in live data pipelines — data corruption, schema drift, backpressure, and failures that only show up at scale.

  • Work with modeling/research partners to understand data quality, format, and structure needs, and translate them into working pipelines.

  • Participate in design and code review across the data infrastructure stack.

Requirements:
  • 2+ years of software engineering experience, with some exposure to data pipelines, data engineering, or backend systems.

  • Strong programming fundamentals in Python, with the ability to write performant, production-grade data processing code.

  • Experience with at least one distributed data processing framework (e.g., Spark, Ray, Dask, or Flink), or strong fundamentals and willingness to ramp up quickly.

  • Familiarity with data storage concepts — object storage, data lake table formats, and warehouse vs. lake tradeoffs.

  • Bias for ownership: you've taken features or systems from prototype to production.

  • Clear communicator who collaborates well with research/modeling partners and more senior teammates.

  • Experience with workflow orchestration tools (e.g., Airflow, Dagster, Prefect) is a plus.

  • Experience with streaming ingestion for high-volume, near-real-time data is a plus.

  • Experience building automated data quality/curation systems (statistical filtering, anomaly detection, deduplication, or using ML models in an annotation/filtering pipeline) is a plus.

  • Experience with robotics- or embodied-AI-specific data (multi-embodiment datasets, teleoperation/demonstration data, egocentric video) is a plus.

Similar Jobs

22 Days Ago
In-Office
Sunnyvale, CA, USA
207K-275K Annually
Expert/Leader
207K-275K Annually
Expert/Leader
Cloud • Information Technology • Machine Learning
Lead design and implementation of multi-regional data platform and stream-processing architecture. Drive event-driven adoption, mentor engineers, ensure scalability, reliability, security, compliance, and participate in on-call operations.
Top Skills: CockroachdbGoKafkaKubernetesLinuxNatsPythonShell ScriptingTidbYdbYugabyte
6 Hours Ago
In-Office
San Mateo, CA, USA
Senior level
Senior level
Artificial Intelligence • Big Data • Machine Learning • Software
Design and build scalable, fault-tolerant query engines and integrations with modern data lake formats; optimize query execution using vectorized processing, cost-based optimization, and caching; extend or contribute to open-source platforms; collaborate with product, data science, and engineering teams to align solutions with business needs.
Top Skills: Apache IcebergSparkAvroC++Delta LakeDockerHudiJavaLinuxOrcParquetPrestoRustScalaSQLTrino
2 Days Ago
Remote or Hybrid
California, USA
Senior level
Senior level
Artificial Intelligence • Computer Vision • Machine Learning • Robotics
Design, build, and operate large-scale crawling, ingestion, and ETL/ELT pipelines for petabyte-scale image and video datasets. Optimize distributed processing (Ray/Spark/Dask), ensure dataset provenance/versioning, reduce cost/latency, and collaborate with CV engineers and ML scientists. Provide technical leadership and mentor junior engineers.
Top Skills: AirflowAws S3CocoDagsterDaskDockerDvcGcp GcsKubernetesLakefsOpencvPandasPascal VocPillowPrefectPysparkPythonRayScikit-ImageSparkSQLYolo

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account