Bitdeer Group

California
214 Total Employees

Jobs at Bitdeer Group

Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.

Recently posted jobs

4 Days AgoSaved
In-Office
San Jose, CA, USA
Software
Architect and scale high-cardinality observability and telemetry infrastructure for AI cloud platforms. Build Kubernetes exporters, eBPF diagnostics, hardware and network monitoring, dashboards, alerting, and GPU/network utilization billing pipelines. Establish telemetry standards for AI workloads, support proactive remediation, collaborate with GPU and scheduling teams, lead architecture reviews, and mentor engineers. The role requires expertise in Go, Kubernetes, Prometheus/OpenTelemetry, Linux performance tuning, eBPF, AI hardware metrics, and large-scale cloud or HPC telemetry systems.
4 Days AgoSaved
In-Office
San Jose, CA, USA
Software
Architects and operates high-performance storage infrastructure for AI training and inference. Responsibilities include developing Kubernetes CSI drivers, integrating GPUDirect Storage, optimizing NVMe caching and I/O performance, supporting RDMA and high-speed networking, monitoring storage systems, managing quotas and multi-tenancy, and mentoring engineers.
4 Days AgoSaved
In-Office
San Jose, CA, USA
Software
Build and maintain SRE microservices supporting a global GPU infrastructure platform. Implement GitOps, declarative configuration, and CI/CD automation; monitor metrics, logs, and traces; support operational readiness through on-call participation, runbooks, and incident reviews; and write unit, integration, and end-to-end tests. Work with senior engineers to deliver production-ready platform features while maintaining service reliability and infrastructure consistency.
4 Days AgoSaved
In-Office
San Jose, CA, USA
Software
Owns architecture, deployment, operations, reliability, and multi-tenant scheduling policy for production Slurm HPC GPU clusters across bare-metal and virtualized environments. Leads Slinky Kubernetes integration, elastic capacity sharing, containerized workloads, GPU and fabric health monitoring, infrastructure automation, observability, accounting, and billing integration. Provides customer onboarding and escalation support, writes operational documentation, leads incident response, and mentors platform engineers.
4 Days AgoSaved
In-Office
San Jose, CA, USA
Software
Architects and operates the hardware foundation for AI cloud infrastructure, integrating NVIDIA and AMD GPUs with Kubernetes. Configures high-performance networking using RDMA, SR-IOV, RoCEv2, and InfiniBand; automates GPU and NIC remediation; manages GPU slicing, kernel and driver tuning, bare-metal provisioning, and firmware updates. Investigates cross-layer performance issues, supports topology-aware scheduling and storage, establishes operational standards, and mentors engineers.
4 Days AgoSaved
In-Office
San Jose, CA, USA
Software
Lead the architecture and implementation of AI workload scheduling and orchestration for a large-scale NeoCloud platform. Build batch and gang scheduling, admission control, queueing, topology-aware placement, GPU sharing, and multi-tenancy capabilities using Kubernetes technologies. Optimize NVLink and InfiniBand utilization, integrate scheduling with hardware and storage systems, improve reliability in HPC environments, resolve contention and deadlocks, and mentor engineers while guiding architectural decisions.
4 Days AgoSaved
In-Office
San Jose, CA, USA
Software
Architect and build a highly available Kubernetes control plane for an AI-native cloud platform managing thousands of clusters. Develop Operators and CRDs, optimize etcd, implement multi-tenant security, enable zero-downtime upgrades and federation, and integrate with GPU orchestration systems. Lead distributed systems architecture, reliability engineering, design reviews, and team mentorship while preventing control-plane failures from affecting customer workloads.
11 Days AgoSaved
In-Office
San Jose, CA, USA
Software
Lead architecture, design, and evolution of a global multi-region cloud SRE platform for GPU/AI compute. Author and maintain platform architecture, enforce design invariants, review framework changes, run plugin framework, decide tier placements, coordinate with cloud teams and security, produce pre-flight designs, and shepherd implementations through engineering squads.
11 Days AgoSaved
In-Office
San Jose, CA, USA
Software
Lead design and implement a global public cloud SRE platform for AI and compute workloads. Own architecture and production engineering for observability, cluster health, remediation, lifecycle, secrets, CI/CD, backup/DR, and automation. Collaborate with cross-functional teams to build scalable, reliable multi-region services and run them in production (on-call).
11 Days AgoSaved
In-Office
San Jose, CA, USA
Software
Design, deploy, and operate production Kubernetes control planes for large GPU clusters. Implement GPU-specific scheduling, CRDs, multi-tenant isolation, BMaaS provisioning, Terraform-based IaC, monitoring, SLI/SLOs, and automated remediation workflows to enable autonomous AIOps-driven recovery and tenant self-service.
11 Days AgoSaved
In-Office
San Jose, CA, USA
Software
Operate and tune InfiniBand and RoCEv2 fabrics for large GPU clusters (100–10,000 GPUs). Monitor UFM and RDMA telemetry, diagnose link/congestion issues, manage firmware and optics lifecycle, collaborate with vendor support, and feed labeled incidents/metrics into AIOps to enable predictive remediation and runbook automation.
11 Days AgoSaved
In-Office
San Jose, CA, USA
Software
Front-line SRE covering 8AM–8PM PST monitoring GPU clusters, networking, storage, and sensors. Execute runbooks, perform hardware triage and physical DC tasks (rack, cable, swap), collect diagnostics for escalation, manage tickets (ServiceNow/Jira), update runbooks, and tag incidents to train the AIOps platform.
11 Days AgoSaved
In-Office
San Jose, CA, USA
Software
Deploy, operate, and optimize high-performance parallel/distributed storage systems for AI training and inference; implement multi-tenant isolation and GPU Direct data paths; instrument telemetry for storage-fault prediction; convert incidents into automated runbook-as-code; plan capacity, firmware, migrations, and DR for GPU clusters.
13 Days AgoSaved
In-Office
San Jose, CA, USA
Software
Design, implement, and maintain CI/CD and MLOps pipelines, provision and scale cloud-native AI infrastructure (Kubernetes, GPU clusters), enforce IaC practices, implement observability and security frameworks, lead incident response and cross-functional platform engineering to ensure high availability and compliance.
13 Days AgoSaved
In-Office
San Jose, CA, USA
Software
Lead end-to-end architecture of AI data center networks, DCI, and global backbone for large GPU clusters. Design underlay and overlay networks, IP/VLAN/VXLAN planning, congestion control (PFC/ECN/INT), and optical transport. Produce HLD/LLD, topology diagrams, standards, and SOPs; engage vendors and drive risk mitigation and network improvements.
15 Days AgoSaved
In-Office
San Jose, CA, USA
Software
Manage and strategize patents, collaborate with R&D on innovations, handle IP disputes, and drive compliance and training efforts.
17 Days AgoSaved
In-Office
San Jose, CA, USA
Software
Lead identification, application, and management of US federal, state, and local tax abatement, exemption, and incentive programs (85%). Support US tax advisory, compliance, and liaison with tax authorities (15%), model financial impact, maintain centralized tracking, manage renewals, and coordinate with stakeholders and external advisors to ensure compliance and optimize tax savings.