Design, build, and operate reliable, high-performance inference platforms for large machine learning models in production. Responsibilities include request routing, batching, caching, autoscaling, GPU utilization, observability, distributed systems, performance engineering, capacity planning, incident response, and cost-efficient serving across diverse workloads.
Machine Learning Infrastructure Engineer – Remote
Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.
This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential.
Job Title: Machine Learning Infrastructure Engineer
Location: 100% Remote (U.S.)
Position Type: Full-time, Direct W2
Salary Range: $105,000–$143,000 Annually
Experience Required: 6+ years
Sponsorship: U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position.
Job Summary
We are seeking a Machine Learning Infrastructure Engineer to design, build, and operate high-performance, highly reliable inference platforms for serving large machine learning models in production. The role focuses on the systems engineering side of AI deployment, including request routing, batching, caching, autoscaling, GPU utilization, and end-to-end observability across diverse model workloads. The ideal candidate brings strong distributed systems and performance engineering expertise, has shipped serving systems at scale, and understands the trade-offs between latency, throughput, cost, and quality in ML serving.
Required Qualifications
Would you like to know more about this opportunity? For immediate consideration, please send your resume to [email protected].
Bright Vision Technologies is an Equal Opportunity Employer.
Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.
This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential.
Job Title: Machine Learning Infrastructure Engineer
Location: 100% Remote (U.S.)
Position Type: Full-time, Direct W2
Salary Range: $105,000–$143,000 Annually
Experience Required: 6+ years
Sponsorship: U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position.
Job Summary
We are seeking a Machine Learning Infrastructure Engineer to design, build, and operate high-performance, highly reliable inference platforms for serving large machine learning models in production. The role focuses on the systems engineering side of AI deployment, including request routing, batching, caching, autoscaling, GPU utilization, and end-to-end observability across diverse model workloads. The ideal candidate brings strong distributed systems and performance engineering expertise, has shipped serving systems at scale, and understands the trade-offs between latency, throughput, cost, and quality in ML serving.
Required Qualifications
- Bachelor’s or Master’s degree in Computer Science or a related field.
- Six or more years of experience in distributed systems, infrastructure, or ML platform engineering.
- Strong proficiency in Python and a systems language such as Go, Rust, or C++.
- Deep experience operating high-throughput, low-latency services in production.
- Hands-on experience with LLM or large model inference frameworks such as vLLM or TensorRT-LLM.
- Strong understanding of GPU architecture, memory hierarchies, and accelerator utilization.
- Familiarity with Kubernetes, autoscaling, and modern cloud platforms.
- Experience with observability stacks including metrics, tracing, and structured logging.
- Solid grounding in performance engineering and capacity planning.
- Strong communication and incident response skills.
- Open-source contributions to model serving infrastructure.
- Experience with multi-region or globally distributed AI serving.
- Familiarity with model quantization, distillation, and compression techniques.
- Exposure to FinOps for AI workloads and cost-efficient serving design.
- Experience supporting external-facing AI APIs at scale.
Would you like to know more about this opportunity? For immediate consideration, please send your resume to [email protected].
Bright Vision Technologies is an Equal Opportunity Employer.
Similar Jobs
Financial Services
Designs, builds, and operates secure, scalable cloud and GPU infrastructure platforms for enterprise AI/ML workloads. Leads architecture, production coding, Kubernetes and container operations, CI/CD, infrastructure automation, performance optimization, and reliability efforts. Partners with AI/ML and platform teams to support distributed multi-GPU training and inference. Provides technical leadership while advancing responsible AI-assisted engineering, secure SDLC practices, automation, and operational excellence.
Top Skills:
BcmC#Ci/CdCloud InfrastructureCudaDistributed SystemsDockerGoGpu InfrastructureJavaKubernetesLinuxMicroservicesMlflowNvidia DcgmNvidia DriversPythonRay.IoSlurm
Artificial Intelligence • Cloud • Machine Learning • Mobile • Software • Virtual Reality • App development
Build and optimize large-scale machine learning infrastructure for content retrieval and recommendation. Responsibilities include developing feature generation and serving pipelines, high-performance inference systems, cloud-based training and evaluation infrastructure, and data management systems. The role partners with ML engineers to deploy models, improve reliability and efficiency, and operate highly available distributed systems using Java, Go, C++, and Python.
Top Skills:
C++Caffe2FlinkGoJavaPythonPyTorchRayScikit-LearnSparkSpark MlTensorFlow
Automotive • Big Data • Information Technology • Robotics • Software • Transportation • Manufacturing
Design and build scalable, reliable ML training infrastructure for advanced AI research and model development. Optimize distributed training performance, resource utilization, data loading, and costs across heterogeneous GPU and cloud environments. Improve platform observability, debuggability, operational excellence, and user experience. Collaborate with ML engineers, research scientists, and cross-functional partners to integrate new technologies supporting intelligent driving systems.
Top Skills:
AWSAzureDistributed ComputingFsdpGCPGpu ComputingPipeline ParallelismPythonPyTorchTensorFlow
What you need to know about the San Francisco Tech Scene
San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.
Key Facts About San Francisco Tech
- Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
- Major Tech Employers: Google, Apple, Salesforce, Meta
- Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
- Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
- Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
- Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine



