Nebius Logo

Nebius

Site Reliability Engineer

Reposted 2 Months Ago
Remote
Hiring Remotely in United States
100K-140K Annually
Mid level
Remote
Hiring Remotely in United States
100K-140K Annually
Mid level
The Linux Systems Administrator will maintain and troubleshoot Linux systems, support network services, and work on systems integration while collaborating with infrastructure teams.
The summary above was generated by AI

About Nebius:

Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure.

Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI.

Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D.

 

The role

Nebius is looking for a Site Reliability Engineer in Hardware Infrastructure team. This is a remote position for the United States.

Hardware Infrastructure team designs, develops and supports systems involved in the data-centers lifecycle:

  • Serving functional and load testing system.
  • Monitoring of engineering equipment located in our data centers (power supply, air and water cooling, etc.)
  • Monitoring of IT equipment: racks, servers, JBODs, JBOGs, power shelves, network devices, etc.
  • Asset tracking.
  • Hardware repairs tasks tracking.
  • Server production.

In this position, your responsibility will be to:

  • Ensure fault-tolerance, scale and uninterrupted operations for our services.
  • Use cutting-edge technology to solve a variety of infrastructure problems.
  • Implement and improve CI/CD processes.

We expect you to have: 

  • Proficiency in Linux systems, with expertise in Python and Bash scripting for automation.
  • Demonstrated ability to troubleshoot complex system issues, including hardware, software and networking problems.
  • Strong analytical and problem-solving skills, with a focus on optimizing system performance.
  • Working proficiency in English.

It would be an added bonus if you had:

  • Desire to be involved in backend development.
  • Experience designing, developing and running high-load distributed systems.

Working conditions:

  • Primarily remote 
  • Occasional travel to data centers required, especially if not located near one
  • Collaboration with globally distributed engineering and operations teams

Key employee benefits:

  • Health insurance: 100% company-paid medical, dental, and vision coverage for employees and families
  • 401(k) plan: up to 4% company match with immediate vesting
  • Parental leave: 20 weeks paid for primary caregivers, 12 weeks for secondary caregivers
  • Remote work reimbursement: up to $85/month for mobile and internet
  • Disability & life insurance: company-paid short-term, long-term, and life insurance coverage

Compensation

  • We offer competitive salaries, ranging from $130k- $180k base + quarterly performance bonuses.

Join Nebius and help operate the systems that power next-generation AI
infrastructure.

Benefits & Perks:

  • Competitive compensation
  • Career growth and learning opportunities
  • Flexibility and ownership
  • Collaborative and innovative culture
  • Opportunity to work on impactful AI projects
  • International environment and talented teams

What's it like to work at Nebius:

Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI 

Equal Opportunity Statement:

Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law.

Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. 

If you need accommodations during the application process, please let us know.

Similar Jobs

4 Days Ago
Remote or Hybrid
United States
111K-180K Annually
Senior level
111K-180K Annually
Senior level
Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Leads the architecture, modernization, optimization, reliability, and performance of enterprise IBM z/OS Db2 systems. Provides technical guidance across development and operations teams, automates DDL/DML processes, supports resilience and business continuity testing, resolves Db2 incidents, analyzes performance telemetry, improves SQL and database design, and partners with architects and stakeholders on enterprise technology strategy and hybrid-cloud modernization.
Top Skills: AnsibleCloud IntegrationDdlDevOpsDmlIbm Db2Ibm Z/OsOpenshiftPythonRed Hat AnsibleRmfSmfSQLZlinux
6 Days Ago
In-Office or Remote
92K-164K Annually
Senior level
92K-164K Annually
Senior level
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Designs and operates secure, reliable Azure cloud platforms using Terraform, GitHub Actions, containers, and automation. Responsibilities include CI/CD, observability, incident response, platform security, vulnerability remediation, disaster recovery, infrastructure troubleshooting, and SRE practices. The role supports production workloads, improves reliability and delivery processes, participates in on-call activities, and mentors engineers while partnering across development, security, architecture, and operations teams.
Top Skills: BashCi/CdCloud SecurityDockerGitGithub ActionsGitopsInfrastructure As CodeKubernetesAzureObservabilityPowershellPythonTerraform
6 Days Ago
Remote or Hybrid
United States
111K-180K Annually
Senior level
111K-180K Annually
Senior level
Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Leads architecture, modernization, optimization, and reliability initiatives for mainframe CICS, MQ, and z/OS Connect environments. Provides technical direction across development and operations teams, establishes governance and change processes, tunes performance using telemetry, resolves incidents, and develops modernization roadmaps. Collaborates with stakeholders and enterprise architects to deliver secure, scalable, high-availability solutions while evaluating automation, cloud integration, and AI technologies.
Top Skills: AnsibleCicsCobolDevOpsIbm MqIbm Z/OsOpenshiftPythonRed Hat Ansible Automation PlatformZ/Os ConnectZlinux

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account