Cohere AI Jobs

Site Reliability Engineer, Inference Infrastructure

Cohere AI

Site Reliability Engineer, Inference Infrastructure

Reposted 9 Days Ago

In-Office or Remote

Hiring Remotely in San Francisco, CA, USA

Senior level

In-Office or Remote

Hiring Remotely in San Francisco, CA, USA

Senior level

The Site Reliability Engineer will develop, deploy, and operate AI infrastructure, focusing on high-performance and scalable machine learning systems using Kubernetes and cloud platforms.

The summary above was generated by AI

Who are we?

Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems.

We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that.

We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft.

We are a global technology company co-headquartered in Toronto and San Francisco, with key offices in London, New York City, Montreal, Seoul, Germany and Paris. Join us!

Why this role?

Are you energized by building high-performance, scalable and reliable machine learning systems? Do you want to help define and build the next generation of AI platforms powering advanced NLP applications? We are looking for a Site Reliability Engineer to join the Model Serving team at Cohere. The team is responsible for developing, deploying, and operating the AI platform delivering Cohere's large language models through easy to use API endpoints. In this role, you will work closely with many teams to deploy optimized NLP models to production in low latency, high throughput, and high availability environments. You will also get the opportunity to interface with customers and create customized deployments to meet their specific needs.

As a Site Reliability Engineer you will:

Build self-service systems that automate managing, deploying and operating services.
This includes our custom Kubernetes operators that support language model deployments.
Automate environment observability and resilience. Enable all developers to troubleshoot and resolve problems.
Take steps required to ensure we hit defined SLOs, including participation in an on-call rotation.
Build strong relationships with internal developers and influence the Infrastructure team’s roadmap based on their feedback.
Develop our team through knowledge sharing and an active review process.

You may be a good fit if you have:

5+ years of engineering experience running production infrastructure at a large scale
Experience designing large, highly available distributed systems with Kubernetes, and GPU workloads on those clusters
Experience with Kubernetes dev and production coding and support
Experience with GCP, Azure, AWS, OCI, multi-cloud on-prem / hybrid serving
Experience in designing, deploying, supporting, and troubleshooting in complex Linux-based computing environments
Experience in compute/storage/network resource and cost management
Excellent collaboration and troubleshooting skills to build mission-critical systems, and ensure smooth operations and efficient teamwork
The grit and adaptability to solve complex technical challenges that evolve day to day
Familiarity with computational characteristics of accelerators (GPUs, TPUs, and/or custom accelerators), especially how they influence latency and throughput of inference.
Strong understanding or working experience with distributed systems.
Experience in Golang, C++ or other languages designed for high-performance scalable servers).

Full-Time Employees at Cohere enjoy these Perks:

A weekly lunch stipend of $75/£75 or equivalent in your local currency for lunch.
Full health and dental benefits, including a separate budget for mental health.
RRSP matching, 401K, Pension Scheme.
100% Parental Leave top-up for up to 6 months, for either parent.
Annual enrichment benefits:
Arts & culture, fitness/wellness, quality time, and a workspace improvement credit.
Education & learning stipend for conferences, courses, and coaching.

6 weeks of paid vacation (30 working days!)
Budget for traveling to other offices if you are remote, plus an annual company offsite.

How and Where We Work:

Cohere is remote-friendly. We have offices in Toronto, San Francisco, New York City, London, Paris, Montreal, and more coming soon.
For those in the office: a daily lunch program, plenty of snacks, and regular community and social events.
For those not near an office: a co-working benefit so you can work alongside others in your city.
Everyone receives a $500 home office stipend to set up your workspace properly.

If any of the above doesn’t line up exactly with your experience, we still encourage you to apply.

We strive to create an inclusive work environment for all; we welcome applicants from all backgrounds and are committed to providing equal opportunities. Should you require any accommodations during the recruitment process, please submit an Accommodations Request Form, and we will work together to meet your needs.

We may use AI-enabled tools to screen and assess applicants against the criteria for this position. This helps our recruiters identify potentially qualified candidates, but it doesn't limit the applications our recruiters may review or consider.

San Francisco, California, United States

Similar Jobs

ServiceNow

Sr. Partner Manager

3 Hours Ago

Remote or Hybrid

Senior level

Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation

Manage and grow ServiceNow partner ecosystem across Canada through partner business planning, enablement, governance, reporting, coaching, and joint GTM to drive partner revenue and program maturity. Conduct reviews, remediation, and cross-functional alignment while supporting partner portal operations and enablement programs.

Top Skills: Ai-Powered ToolsServicenow

Applied Systems

Senior Compensation Analyst

5 Hours Ago

Remote or Hybrid

100K-135K Annually

Senior level

100K-135K Annually

Senior level

Cloud • Insurance • Payments • Software • Business Intelligence • App development • Big Data Analytics

Lead development and administration of compensation frameworks, manage the annual compensation planning cycle, perform market benchmarking and pay equity analyses, administer equity programs, own compensation system administration and automation, support international pay benchmarking (including India), translate findings into cost/benefit recommendations, ensure compliance with compensation laws, and mentor junior team members and HR partners.

Top Skills: AIAonBettercompCulpepperData AnalyticsExcelHrisPequityRadfordUkgWtw

Atlassian

Account Executive

7 Hours Ago

Remote

95K-123K Annually

Mid level

95K-123K Annually

Mid level

Cloud • Information Technology • Productivity • Security • Software • App development • Automation

Manage a ~40-account mid-market portfolio to drive net-new growth and expansion. Own full sales cycle, develop account/territory plans, lead cross-functional deal teams, build C‑suite relationships, qualify and close complex multithreaded deals using outcome-based selling, forecast pipeline, and collaborate with channel, product, and customer success. Occasional travel for customer meetings and team events.

Top Skills: CloudConfluenceCRMJira Service ManagementJira SoftwareSaaS

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
Major Tech Employers: Google, Apple, Salesforce, Meta
Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Cohere AI

Site Reliability Engineer, Inference Infrastructure

Cohere AI San Francisco, California, USA Office

Similar Jobs

Sr. Partner Manager

Senior Compensation Analyst

Account Executive

What you need to know about the San Francisco Tech Scene

Key Facts About San Francisco Tech