Bot Auto Logo

Bot Auto

GPU Engineer

Posted One Month Ago
In-Office or Remote
6 Locations
Mid level
In-Office or Remote
6 Locations
Mid level
Optimize end-to-end GPU performance for real-time autonomous driving workloads: develop and optimize parallel GPU algorithms and inference pipelines, profile and eliminate bottlenecks across computation and memory, debug GPU software for latency/throughput improvements, and collaborate on onboard GPU software architectures for embedded platforms.
The summary above was generated by AI
Company Introduction

At Bot Auto, we are revolutionizing the transportation of goods with our cutting-edge autonomous trucks, enhancing the quality of life for communities around the globe. With the agility of a start-up and the wisdom of seasoned experts, Bot Auto boasts a team that has achieved numerous world-firsts and unparalleled innovations. United by a shared vision, we create miracles and propel the future of transportation. Join us and transform your dreams into reality.

You would collaborate with software engineers, AI researchers, and hardware specialists to develop high-performance solutions that meet the stringent requirements of autonomous driving applications. This is an exciting opportunity to work on next-generation transportation technology and make a meaningful impact on the future of mobility.

Key Responsibilities
  • Optimize end-to-end GPU performance for real-time autonomous driving workloads, including sensor processing (e.g., camera, LiDAR) and neural network inference.
  • Develop and optimize parallel computing algorithms and GPU-accelerated components using technologies such as CUDA.
  • Collaborate with cross-functional teams to design and improve onboard GPU software architectures that meet the computational requirements of perception, planning, and control modules.
  • Profile and analyze bottlenecks across GPU computation, memory access, data movement, synchronization, and CPU–GPU interaction.
  • Debug and optimize GPU-based software to improve latency, throughput, resource utilization, and runtime stability on embedded platforms.
Qualifications:

Required:

  • Bachelor’s or Master’s degree in Computer Science, Electrical Engineering, or a related field.
  • Strong knowledge of parallel computing principles, GPU architecture, memory hierarchy, and performance optimization techniques.
  • Experience profiling GPU applications using tools such as NVIDIA Nsight Systems, Nsight Compute, or equivalent tools.
  • Experience deploying or optimizing neural network inference workloads using technologies such as PyTorch, ONNX, and TensorRT.
  • Experience with real-time embedded systems and handling large data streams from sensors (camera, LiDAR, radar).
  • Strong proficiency in C/C++ and Python.

Preferred:

  • 3+ years of experience in GPU programming and optimization (e.g., CUDA, OpenCL, Vulkan).
  • Experience with NVIDIA Jetson Thor, NVIDIA DRIVE Thor, or similar embedded GPU platforms.
  • Experience with model quantization, including FP8 and NVFP4.
  • Experience managing concurrent GPU workloads and resource isolation using technologies such as NVIDIA Multi-Process Service (MPS), Multi-Instance GPU (MIG), or other related technologies.
  • Experience with GPU-accelerated sensor data compression, including camera, LiDAR, or other onboard sensor data.

Similar Jobs

Yesterday
Remote
United States
90-120 Hourly
Junior
90-120 Hourly
Junior
Artificial Intelligence • HR Tech • Professional Services • Software
Build and evaluate MLOps and ML systems tasks for frontier AI training data. Responsibilities include designing technical challenges, writing solutions and evaluation rubrics, profiling and optimizing GPU workloads, debugging distributed systems, improving model performance, and supporting high-throughput LLM inference. The role requires production experience with ML infrastructure, serving systems, GPU accelerators, JAX or PyTorch, and strong technical communication.
Top Skills: A100B200Continuous BatchingCudaDdpDeepspeedFsdpH100JaxKinetoKv CacheMegatronNsightPaged AttentionPallasPyTorchRay ServeSglangTensorrt-LlmTorch.ProfilerTpuTritonVllmXla
One Month Ago
Remote
United States
Senior level
Senior level
Artificial Intelligence • Hardware • Software • Semiconductor
Build, productionize, and optimize a GPU-based inference stack combining GPU prefill with Cerebras decode. Implement and operate model-serving APIs, vLLM/PyTorch/ROCm runtimes, deployment and reliability practices, performance profiling and optimization, cross-layer debugging, numerical validation, and benchmarking/infrastructure for production inference at scale.
Top Skills: C++Ci/CdContainersCudaKubernetesLinuxPythonPyTorchRdmaRocmSglangTensorrt-LlmTriton Inference ServerVllm
15 Days Ago
Remote or Hybrid
United States
285K-340K Annually
Senior level
285K-340K Annually
Senior level
Artificial Intelligence • Machine Learning • Natural Language Processing • Software • Generative AI
Build and scale ML-optimized HPC infrastructure, manage Kubernetes-based GPU/TPU superclusters, optimize for AI/ML training, and mentor teams while innovating in ML infrastructure.
Top Skills: GoGpuJaxKubernetesLinuxNcclPythonPyTorchRdmaTensorFlowTpu

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account