Bright Vision Technologies Logo

Bright Vision Technologies

Machine Learning Infrastructure Engineer

Posted 5 Days Ago
Be an Early Applicant
In-Office
Santa Clara, CA, USA
105K-143K Annually
Senior level
In-Office
Santa Clara, CA, USA
105K-143K Annually
Senior level
Design, build, and operate high-performance inference platforms for large ML models. Focus on request routing, batching, caching, autoscaling, GPU utilization, observability, performance engineering, and production reliability for diverse model workloads.
The summary above was generated by AI
Machine Learning Infrastructure Engineer – Remote
Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.
This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential.
Job Title: Machine Learning Infrastructure Engineer
Location: 100% Remote (U.S.)
Position Type: Full-time, Direct W2
Salary Range: $105,000–$143,000 Annually
Experience Required: 6+ years
Sponsorship: U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position.
Job Summary
We are seeking a Machine Learning Infrastructure Engineer to design, build, and operate high-performance, highly reliable inference platforms for serving large machine learning models in production. The role focuses on the systems engineering side of AI deployment, including request routing, batching, caching, autoscaling, GPU utilization, and end-to-end observability across diverse model workloads. The ideal candidate brings strong distributed systems and performance engineering expertise, has shipped serving systems at scale, and understands the trade-offs between latency, throughput, cost, and quality in ML serving.
Required Qualifications
  • Bachelor’s or Master’s degree in Computer Science or a related field.
  • Six or more years of experience in distributed systems, infrastructure, or ML platform engineering.
  • Strong proficiency in Python and a systems language such as Go, Rust, or C++.
  • Deep experience operating high-throughput, low-latency services in production.
  • Hands-on experience with LLM or large model inference frameworks such as vLLM or TensorRT-LLM.
  • Strong understanding of GPU architecture, memory hierarchies, and accelerator utilization.
  • Familiarity with Kubernetes, autoscaling, and modern cloud platforms.
  • Experience with observability stacks including metrics, tracing, and structured logging.
  • Solid grounding in performance engineering and capacity planning.
  • Strong communication and incident response skills.
Preferred Qualifications
  • Open-source contributions to model serving infrastructure.
  • Experience with multi-region or globally distributed AI serving.
  • Familiarity with model quantization, distillation, and compression techniques.
  • Exposure to FinOps for AI workloads and cost-efficient serving design.
  • Experience supporting external-facing AI APIs at scale.
How to Apply
Would you like to know more about this opportunity? For immediate consideration, please send your resume to [email protected].
Bright Vision Technologies is an Equal Opportunity Employer.

Similar Jobs

3 Days Ago
Hybrid
190K-246K Annually
Senior level
190K-246K Annually
Senior level
Mobile • Social Media
Design, build, and maintain scalable ML infrastructure and data pipelines for training, deployment, monitoring, and evaluation of large-scale models (including LLMs). Build APIs and platform services, optimize compute/storage, implement CI/CD and GitOps, and support recommendation/moderation systems. Mentor engineers, lead cross-team initiatives, run A/B testing and observability, and research emerging ML infrastructure technologies.
Top Skills: Amazon EcsAmazon EksApache FlinkApache KafkaSparkAWSAzureBuildkiteDatabricksDelta LakeDockerDynamoDBGCPGitopsGoGrafanaGrafana MimirGraphQLGrpcHelmJavaJenkinsPrometheusPythonRay ServeRedisRestScaffoldScalaTerraformTerragruntTritonValkey
4 Days Ago
Hybrid
Palo Alto, CA, USA
235K-414K Annually
Expert/Leader
235K-414K Annually
Expert/Leader
Artificial Intelligence • Cloud • Machine Learning • Mobile • Software • Virtual Reality • App development
Lead technical strategy, architecture, and implementation for ML inference platform services. Design and scale distributed, high-throughput inference systems, collaborate across teams, drive availability, scalability, operational excellence, cost management, and provide company-wide technical direction and mentorship.
Top Skills: Distributed SystemsGpuKubernetesLlm InferenceMl Inference PlatformPyTorchRpcTensorFlow
7 Days Ago
In-Office
San Francisco, CA, USA
224K-294K Annually
Senior level
224K-294K Annually
Senior level
Artificial Intelligence • Software
Turn research prototypes into scalable, fault-tolerant ML and physics infrastructure for drug discovery. Improve GPU utilization, distributed execution, reproducibility, orchestration, monitoring, and packaging so scientists can run reliable, multi-cluster workflows and agent-usable pipelines.
Top Skills: ArgoContainersCudaCuda-Aware WorkflowsDockerFlyteGpu ComputingJaxKubernetesLinuxPythonPyTorchRaySlurm

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account