Genesis Logo

Genesis

Inference

Posted 7 Hours Ago
Be an Early Applicant
In-Office or Remote
6 Locations
Senior level
In-Office or Remote
6 Locations
Senior level
Build and optimize low-latency on-device and distributed GPU inference pipelines. Implement and tune low-level CUDA/Triton kernels, optimize for throughput and latency, and develop monitoring/debugging tools to ensure reliable, deterministic inference in robotics and cluster environments.
The summary above was generated by AI
What You’ll Do
  • Build low-latency inference pipelines for on-device deployment, enabling real-time next-token and diffusion-based control loops in robotics

  • Design and optimize distributed inference systems on GPU clusters, pushing throughput with large-batch serving and efficient resource utilization

  • Implement efficient low-level code (CUDA, Triton, custom kernels) and integrate it seamlessly into high-level frameworks

  • Optimize workloads for both throughput (batching, scheduling, quantization) and latency (caching, memory management, graph compilation)

  • Develop monitoring and debugging tools to guarantee reliability, determinism, and rapid diagnosis of regressions across both stacks

What You’ll Bring
  • Deep experience in distributed systems, ML infrastructure, or high-performance serving (8+ years)

  • Production-grade expertise in Python, with strong background in systems languages (C++/Rust/Go)

  • Low-level performance mastery: CUDA, Triton, kernel optimization, quantization, memory and compute scheduling

  • Proven track record scaling inference workloads in both throughput-oriented cluster environments and latency-critical on-device deployments

  • System-level mindset with a history of tuning hardware–software interactions for maximum efficiency, throughput, and responsiveness

Similar Jobs

17 Days Ago
In-Office or Remote
6 Locations
Senior level
Senior level
Artificial Intelligence • Big Data • Machine Learning
Own and scale WEKA's global partner ecosystem across OEMs and ISVs: manage relationships, define partner GTM and program structure, enable field sales with playbooks and co-sell motions, orchestrate cross-functional certification and co-marketing, hire and lead regional partner managers, and own partner-influenced pipeline and revenue targets.
Top Skills: Ai/MlAnyscaleGpuHigh-Performance StorageInferactInferenceMlopsNeuralmeshOrchestrationSpectro Cloud
Senior level
Artificial Intelligence • Hardware • Software • Semiconductor
Design, build, and operate CI/CD, Kubernetes-based platforms, deployment automation, and observability for engineering workflows. Improve reliability, performance, and scalability across cloud and on-prem environments, debug cross-boundary failures, perform root-cause analysis, and deliver durable platform software and self-service tooling.
Top Skills: Argo CdArtifact RepositoriesAWSCi/CdContainerized EnvironmentsCustom ResourcesHelmKubernetesKubernetes OperatorsLinuxMtlsPackage RegistriesPythonShellTerraformTls
6 Hours Ago
Remote or Hybrid
US
Mid level
Mid level
Artificial Intelligence • Hardware • Software • Semiconductor
Design, implement, and maintain Python frameworks and services that orchestrate distributed engineering workflows across machines and clusters. Build scheduling, execution, resource management, failure recovery, and test infrastructure. Define APIs and abstractions, reason about concurrency and distributed-systems behavior, debug complex multi-system issues, write automated tests and documentation, and partner with platform, CI, release, QA, and product teams to deliver scalable infrastructure.
Top Skills: AsyncioBuild SystemsCi SystemsCluster SchedulersConcurrent.FuturesContainersKubernetesMultiprocessingPytestPythonRelease InfrastructureRemote Execution Systems

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account