Voxel Logo

Voxel

Senior Software Engineer, ML Infrastructure

Reposted 4 Days Ago
Be an Early Applicant
Hybrid
San Francisco, CA, USA
200K-240K Annually
Senior level
Hybrid
San Francisco, CA, USA
200K-240K Annually
Senior level
The role involves managing ML infrastructure, building scalable data pipelines, operating training frameworks, leading projects, and ensuring best DevOps practices for machine learning.
The summary above was generated by AI
Who We Are

Voxel is building the future of Computer Vision and Machine Learning for operations, risk, and safety. We use computer vision and AI to enable existing security cameras to automatically detect hazards and high-risk activities, keep people safe and drive operational efficiencies. Our technology addresses the key cost drivers for workers’ compensation, general liability, and property damage, which cost US employers over $500 billion annually. Our customers include Fortune 500 companies across grocery, retail, manufacturing, food and beverage, logistics, and pharmaceutical distribution. We’ve passed $10M ARR with strong expansion revenue. Based in SF, backed by industry-leading VCs.

 
About the Role

Voxel’s perception system is the technical core of everything we ship. Our models detect human activity, equipment interactions, environmental hazards, and operational state in real time across thousands of cameras in manufacturing, logistics, retail, and pharmaceutical environments. Safety was our wedge; it proved our platform works. Now customers are pulling us into operations: equipment utilization, workflow compliance, process efficiency. Every new use case runs through the perception team.

We're hiring a strong software engineer to own the ML Infrastructure that powers how Voxel trains and ships vision models. You’ll build systems that let our applied ML team train multiple models concurrently, manage experiments and ship optimized models to production. You'll set technical direction, write code, make architecture calls, and partner closely with applied CV, ML Data and Platform engineers.

What You'll Do
  • Build and maintain training infrastructure that lets the applied ML team train multiple models concurrently, manage experiments, and iterate quickly on new architectures.

  • Own the train-to-deploy handoff - export trained models to optimized inference formats (TensorRT, ONNX), quantify accuracy and latency impact, and partner with Platform on production deployment.

  • Establish ML experiment tracking and lifecycle management - pick the right tools (Weights & Biases, MLflow, ClearML, or similar) so researchers can run, compare, and reproduce experiments efficiently.

  • Establish DevOps-for-ML best practices on AWS (IaC, CI/CD, observability, cost monitoring) so researchers can iterate quickly and safely.

  • Understand the infra needs of applied ML/CV engineers and design scalable solutions that support model development.

What We're Looking For
  • 4+ years of experience building and shipping large scale software solutions.

  • Hands-on experience building ML training pipelines in PyTorch.

  • Hands-on experience with ML experiment tracking and lifecycle tools (Weights & Biases, MLflow, ClearML, or similar).

  • Experience with AWS (S3, EC2, EKS, or similar) for ML workloads.

  • Strong Python. Write performant code that scales well in production environments.

  • Track record of owning infrastructure end-to-end: scoping, building, shipping, and improving systems that internal teams depend on.

  • Bias toward shipping. You'd rather ship something good this week than something perfect next quarter.

  • Strong communication skills.

Nice to Have
  • Experience with modern ML orchestration tools (Ray, Sematic, Flyte, Metaflow, Prefect, or similar)

  • Familiarity with GPU performance profiling and optimization (Nsight, PyTorch profiler, or similar)

  • Background in computer vision model training

Compensation & Benefits
  • Equity through Voxel’s Equity Incentive Plan

  • Total compensation includes base salary, annual bonus, and equity

  • Comprehensive health, dental, and vision insurance

  • Competitive paid parental leave

  • Unlimited PTO and flexible work arrangements

  • Daily meals in-office, team events, annual company onsite

Similar Jobs

7 Hours Ago
In-Office
San Francisco, CA, USA
170K-190K Annually
Senior level
170K-190K Annually
Senior level
Artificial Intelligence • Logistics • Robotics • Software
Design, build, and operate large-scale ML infrastructure and GPU compute clusters for computer vision and multi-modal models. Own end-to-end pipelines from data ingestion and training to low-latency cloud deployment, orchestration, monitoring, and model performance evaluation. Collaborate with research and product teams to productionize models across warehouse environments.
Top Skills: AirflowC++CudaDeepspeedFlyteGpuGrpcKafkaNvidia Triton Inference ServerPythonPyTorchQuantizationRos2TemporalTensorFlowTensorflow ServingTensorrtTorchserve
21 Days Ago
In-Office
San Francisco, CA, USA
176K-220K Annually
Senior level
176K-220K Annually
Senior level
Edtech • Enterprise Web • HR Tech • Software
Build and operate shared ML/AI infrastructure: data pipelines, feature stores, training, model serving and LLM platform work. Scale inference and training (GPU, batching, autoscaling), enable evaluation and post-training workflows, and partner with AI, data science, and product teams to productionize models and improve platform reliability and developer experience.
Top Skills: AirflowApache BeamAutoscalingAWSBatchingBigQueryCi/CdDataflowDockerEmbeddingsFeature StoresGCPGoGpu InferenceKubernetesLlmsModel ServingPythonSparkStreaming PipelinesTerraformTypescript
10 Days Ago
In-Office
Sunnyvale, CA, USA
153K-222K Annually
Senior level
153K-222K Annually
Senior level
Hardware • Industrial
Build and operate end-to-end ML infrastructure: distributed cloud GPU training, datasets, training frameworks, evaluation, deployment, and integrate pipelines into product workflows while collaborating with modeling teams.
Top Skills: Apache AirflowCloud GpuDistributed Gpu TrainingFlyteMl PipelinesNvidia TritonPyTorchTensorFlowTensorflow ServingTorchserve

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account