C-Gen.AI Logo

C-Gen.AI

HPC & MLOps Engineer

Posted Yesterday
Remote or Hybrid
2 Locations
Mid level
Remote or Hybrid
2 Locations
Mid level
Design, deploy, and maintain HPC and MLOps infrastructure across cloud and on-prem clusters. Manage schedulers (Slurm/PBS), optimize MPI/CUDA stacks, storage and networking, automate deployments with Python/Bash, instrument systems, and enable AI training/inference pipelines while collaborating with product and support teams.
The summary above was generated by AI
About C-Gen.AI

At C-Gen.AI, we are pioneering the next generation of AI infrastructure. As a leading technology company, our mission is to deliver world-class AI Infrastructure management solutions that drive innovation, efficiency, and performance at scale. We seek passionate, highly skilled engineers who excel in dynamic environments and are eager to work with cutting-edge technologies to transform the computational landscape.

Role Summary

As an HPC & MLOps Engineer, you will be a foundational member of our Data Center team—integral to the design, deployment, and maintenance of high-performance computing (HPC) and MLOps systems. In this role, you will ensure our clients have access to reliable, scalable, and secure HPC environments. You will collaborate closely with cross-functional teams and work under a dynamic, startup environment that rewards initiative and technical excellence.

Key Responsibilities
  • HPC Infrastructure Support:

    • Configure, deploy, monitor, and maintain C-Gen.AI Cluster solution for a diverse client base.

    • Manage both cloud and on-premise deployments, ensuring optimal job scheduling and resource allocation.

    • Troubleshoot and optimize HPC library stacks—including OpenMPI, CUDA, TensorFlow, and PyTorch—and manage parallel file systems (e.g. Lustre, BeeGFS, Ceph, NFS, or object storage).

  • Cloud & On-Prem Automation:

    • Develop and oversee automated deployments across multiple platforms (Cloud providers or on-premise clusters).

    • Implement best practices for network configuration, security, and cost optimization tailored to HPC needs.

    • Create and maintain Bash/Python scripts that streamline workflows, gather essential metrics, and empower self-service HPC management.

  • MLOps Enablement: Contribute to the development and maintenance of HPC workflows for AI/ML teams, enhancing training and inference pipelines.

  • Potential HPC Development:

    • Collaborate with our in-house HPC experts on advanced projects involving performance tuning, MPI, and GPU parallelization.

    • While not mandatory at present, there is ample opportunity to expand your technical skill set in HPC development over time.

  • Reliability & visibility: Instrument systems, collect metrics, and build dashboards/alerts that enable self-service and rapid incident response.

  • Collaboration: Work closely with product, support, and customer teams; document designs, runbooks, and standards

What You Bring
  • Experience: 3+ years hands-on in HPC operations; strong familiarity with Slurm (or similar batch systems such as PBS Pro/OpenPBS).

  • Software skills: Solid Python (including asyncio/task-oriented patterns) and Bash for automation, tooling, and data handling.

  • Scheduling & process management: Deep understanding of batch queues, job lifecycle, and multi-tenant cluster policies.

  • Linux: Strong understanding of RedHat/Debian linux flavors and system administration.

  • Cloud & hybrid: Proficiency with at least one major cloud (AWS/GCP/Azure/OCI) and interest in others; comfortable bridging cloud and on-premises deployments.

  • Storage & performance: Practical knowledge of Lustre/Ceph/NFS/object storage and the throughput/latency trade-offs common to HPC.

  • ML/HPC fluency: Understanding of AI training/inference workflows, GPU scheduling, drivers, and runtime management.

  • Networking: Confidence with high-performance networking (e.g., RDMA/InfiniBand/RoCE) and compiling Linux modules for network support.

  • Communication: Clear written/spoken English; able to collaborate across time zones and functions.

  • Mindset: Self-directed, detail-oriented, and comfortable in a fast-moving startup.

Nice to Have
  • Python libraries & SDKs: asyncio, aiohttp, cloud SDKs (e.g., boto3, google-cloud, azure-sdk, OCI).

  • Performance tuning: CUDA/NCCL profiling, MPI optimization, kernel/sysctl tuning, GPU/CPU/IO benchmarking.

  • Cost optimization: Experience balancing cost, performance, and reliability—especially in cloud bursting scenarios.

  • C/C++ Experience: Having C/C++ and systems programming experience is a big plus.

Why Join C-Gen.AI?
  • High Impact & Growth: Play a crucial role in a startup with unicorn potential, driving innovation at every level.

  • Ownership & Influence: Enjoy significant autonomy with a direct influence on both technical decisions and team culture.

  • Competitive Compensation: Competitive salary, stock options, and flexible remote work arrangements.

  • Mentorship & Learning: Collaborate directly with our technical CEO and an experienced HPC specialist, ensuring continuous professional growth.

  • Culture of Autonomy: Thrive in an environment that values trust, minimal supervision, and individual initiative.

If you’re excited to build the foundation of large-scale AI, we’d love to hear from you. Apply now and help shape the future of AI infrastructure.

Similar Jobs

6 Hours Ago
Remote
United States
232K-348K Annually
Senior level
232K-348K Annually
Senior level
Artificial Intelligence • Productivity • Software • Automation
As a Sr. Applied AI Engineer at Zapier, you will build and enhance AI platform capabilities, focusing on LLM Ops and ML Ops to support scalable AI development across teams.
Top Skills: Cloud InfrastructureLlm OpsMl OpsPythonTypescript
Yesterday
Remote or Hybrid
Senior level
Senior level
Big Data • Food • Hardware • Machine Learning • Retail • Automation • Manufacturing
Perform FP&A work for SMG&A Overheads across DACH & CEE: collect and structure data, consolidate expenditure, run reconciliations and variance analysis in multiple systems, support planning/forecasting and accruals, track optimization projects, deliver ad hoc analyses, ensure controls and collaborate with global finance teams to improve processes.
Top Skills: AdaptiveCmtExcelFitPowerPointSacSAP
Yesterday
Easy Apply
Remote or Hybrid
Easy Apply
Senior level
Senior level
Artificial Intelligence • Cloud • Information Technology • Machine Learning • Software
The Account Executive will drive new business by selling SaaS solutions to Managed Service Providers in the Nordics, managing the full sales cycle, and achieving revenue targets.
Top Skills: AICRMMachine LearningMeddpiccSaaS

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account