Build and operate large-scale data systems for AI training and evaluation. Design ingestion, transformation, quality, lineage, and high-throughput delivery pipelines for ML workloads. Operate petabyte-scale storage, ensure reproducibility and dataset versioning, and collaborate across teams to optimize data infrastructure for model quality and training efficiency.
Machine Learning Data Engineer – Remote
Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.
This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential.
Job Title: Machine Learning Data Engineer
Location: 100% Remote (U.S.)
Position Type: Full-time, Direct W2
Salary Range: $80,000–$100,000 Annually
Experience Required: 6+ years
Sponsorship: U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position.
Job Summary
We are seeking an Machine Learning Data Engineer to build and operate the large-scale data systems that power modern AI training and evaluation pipelines. The role combines deep data engineering expertise with a strong understanding of AI workloads, focusing on ingestion, transformation, quality assurance, lineage, and high-throughput delivery of data to training jobs across diverse modalities. The ideal candidate has experience operating petabyte-scale data systems, strong software engineering fundamentals, and clear understanding of how data infrastructure choices propagate into model quality and training efficiency.
Required Qualifications
Would you like to know more about this opportunity? For immediate consideration, please send your resume to [email protected].
Bright Vision Technologies is an Equal Opportunity Employer.
Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.
This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential.
Job Title: Machine Learning Data Engineer
Location: 100% Remote (U.S.)
Position Type: Full-time, Direct W2
Salary Range: $80,000–$100,000 Annually
Experience Required: 6+ years
Sponsorship: U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position.
Job Summary
We are seeking an Machine Learning Data Engineer to build and operate the large-scale data systems that power modern AI training and evaluation pipelines. The role combines deep data engineering expertise with a strong understanding of AI workloads, focusing on ingestion, transformation, quality assurance, lineage, and high-throughput delivery of data to training jobs across diverse modalities. The ideal candidate has experience operating petabyte-scale data systems, strong software engineering fundamentals, and clear understanding of how data infrastructure choices propagate into model quality and training efficiency.
Required Qualifications
- Bachelor’s or Master’s degree in Computer Science or a related field.
- Six or more years of data engineering experience, with significant work supporting ML or AI workloads.
- Strong proficiency in Python and at least one JVM or systems language.
- Deep experience with modern data processing frameworks such as Spark, Ray, or Beam.
- Hands-on experience operating petabyte-scale storage and pipeline systems.
- Strong understanding of distributed systems, data modeling, and storage formats.
- Experience with dataset versioning, lineage, and reproducibility for ML workflows.
- Familiarity with high-throughput data loading for accelerator-based training.
- Strong software engineering practices including testing, CI/CD, and code review.
- Excellent communication and cross-functional collaboration skills.
- Experience with multimodal datasets at large scale.
- Familiarity with data quality tooling and dataset evaluation methodology.
- Exposure to privacy-preserving data systems and regulated data handling.
- Open-source contributions to data infrastructure projects.
- Experience supporting frontier model training pipelines.
Would you like to know more about this opportunity? For immediate consideration, please send your resume to [email protected].
Bright Vision Technologies is an Equal Opportunity Employer.
Similar Jobs
Financial Services
Lead design and delivery of reliability, observability, and SRE practices for large-scale data and AI/ML platforms. Define NFRs, SLI/SLOs, incident response, and implement reliable, secure, scalable infrastructure and automation while mentoring teams and driving safe, auditable AI-assisted operations.
Top Skills:
AWSAws GlueCi/CdDatabricksDatadogDockerDynatraceGrafanaKubernetesMapreducePrometheusPythonSparkSplunkTerraform
AdTech • Artificial Intelligence • Big Data • Machine Learning • Marketing Tech • Mobile • Software
Build and maintain scalable, reliable ML data platform systems to support dataset generation, model training, analytics, monitoring, and large-scale inference. Collaborate with ML, software, and infrastructure engineers to design cost-efficient data lake and training infrastructure using vendor and open-source tools to enable next-generation ML models.
Top Skills:
GoPython
Artificial Intelligence • Greentech • Hardware • Internet of Things • Transportation • Cybersecurity • Automation
Build and consolidate data infrastructure for ML training: design automated ETL/ELT pipelines, implement DataOps practices (validation, monitoring, anomaly detection), integrate annotation workflows, manage dataset versioning and storage, and collaborate with ML engineers and operations to produce high-quality, reproducible training datasets.
Top Skills:
Apache AirflowBigQueryCloud StorageDagsterData ValidationDataflowDataprocDvcGoogle Cloud ComposerGreat ExpectationsNumpyPandasPrefectPythonSQLTfxVertex Ai Data Pipelines
What you need to know about the San Francisco Tech Scene
San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.
Key Facts About San Francisco Tech
- Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
- Major Tech Employers: Google, Apple, Salesforce, Meta
- Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
- Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
- Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
- Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine



