Metasys Logo

Metasys

Site Reliability Engineer Internship

Reposted Yesterday
Remote
Hiring Remotely in United States
Internship
Remote
Hiring Remotely in United States
Internship
Internship SRE role responsible for availability, performance, and scalability of an e-commerce supply-chain platform. Tasks include SLO/SLA definition, observability (Prometheus/Grafana/Loki/Tempo/OpenTelemetry), incident response, capacity planning, disaster recovery for PostgreSQL, infrastructure-as-code (Terraform), CI/CD automation, and operational reliability for AI agent services. Mentored by Head of Technology/CTO with potential conversion to full-time based on performance.
The summary above was generated by AI
Overview: Reliability and Operational Excellence

The Site Reliability Engineer (SRE) is responsible for the ultimate stability, performance, and scalability of our entire integrated supply chain e-commerce platform. You will apply software engineering principles to operations, ensuring the high availability and resilience of the customer-facing e-commerce storefront, internal SaaS tools (WMS, OMS), and specialized AI agent services.

Internship Details

Duration: 3 months
Start Date: Immediate
Location: Remote
Stipend: None initially. Based on your first-quarter performance, you may be offered a paid full-time opportunity, or even be absorbed directly by the client as an FTE.

Key Responsibilities & Core Projects

You will be the champion of uptime, performance, and automated operations for systems handling the critical MES → WMS → OMS flow.

  • Availability & SLO Management: Define, implement, and track Service Level Objectives (SLOs) and Service Level Indicators (SLIs) for core business processes and all application layers. Manage the platform's overall Service Level Agreement (SLA).

  • Observability & Alerting: Architect, maintain, and optimize the comprehensive observability stack (Prometheus, Grafana, Loki, Tempo, OpenTelemetry). Develop high-fidelity alerting and ensure distributed tracing across the NestJS modular monolith and associated data stores (PostgreSQL, Redis).

  • Incident Response & Review: Own the incident response workflow, ensuring rapid triage, mitigation, and root cause analysis. Conduct thorough post-incident reviews to drive continuous improvement and eliminate recurring toil.

  • Scalability & Capacity Planning: Optimize auto-scaling policies for all services running on Docker containers. Conduct capacity planning based on business projections, especially for peak e-commerce and manufacturing load.

  • Disaster Recovery (DR): Design, implement, and regularly test Disaster Recovery procedures, including backup and restoration workflows for PostgreSQL 15 using tools like pgBackRest.

  • Automation: Eliminate operational toil through automation, managing infrastructure-as-code (Terraform) and CI/CD pipelines (Makefile).

Required Technologies & Tools

Candidates must possess deep experience in cloud operations, observability, and infrastructure automation:

  • Observability Stack: Prometheus, Grafana, Loki, Tempo, OpenTelemetry (mandatory).

  • Infrastructure & Platform: Terraform, Docker, Traefik, Oracle Cloud Free VMs (or equivalent public cloud).

  • Data & Resilience: PostgreSQL (Deep knowledge), Redis, pgBackRest.

  • Automation: Strong scripting skills (Python/Bash) and experience with CI/CD tools and Makefile.

  • Methodology: Expert knowledge of SRE principles, toil reduction, and error budgeting.

AI Agent Focus

You will be responsible for the operational reliability of the emerging AI layer.

  • Agent Reliability: Implement specialized monitoring and logging for the AI agent services, ensuring LLM integrations and multi-agent systems (built with frameworks like LangChain) meet defined performance and availability SLOs.

  • Resource Optimization: Efficiently manage resource allocation for computationally intensive AI workloads to maintain platform stability and cost-efficiency.

Success Metrics & Career Path

Performance will be measured by:

  • Uptime/Availability: Achieving defined SLAs/SLOs across the platform.

  • MTTR: Reduction in Mean Time To Recover from production incidents.

  • Toil Reduction: Measured percentage reduction in manual, repetitive operational tasks through automation.

Mentorship Structure: Reports to the Head of Technology/CTO, working collaboratively with DevSecOps and development teams to ensure software is designed for reliability.

Similar Jobs

45 Minutes Ago
Remote or Hybrid
USA
91K-203K Annually
Senior level
91K-203K Annually
Senior level
Machine Learning • Payments • Security • Software • Financial Services
Leads complex, multi-workstream product initiatives as a principal Product Owner. Owns product vision, strategy, backlog integrity, prioritization, and Scrum team alignment. Coordinates Product, Engineering, Design, Delivery, and other stakeholders; improves execution discipline, transparency, and outcome tracking. Stabilizes programs with fragmented ownership or elevated delivery risk while ensuring customer and business requirements are addressed. Requires substantial Product Owner or Senior Product Manager experience, preferably in payments or financial services.
Top Skills: Agile DevelopmentData VisualizationScrumUx Design
2 Hours Ago
Remote or Hybrid
MI, USA
Senior level
Senior level
Software • Analytics • Hospitality
Leads revenue recognition, billing, accounts receivable, collections, audit readiness, and financial control operations. Ensures ASC 606 and GAAP compliance, manages billing platforms and ERP integrations, drives automation and AI adoption, develops scalable processes, partners cross-functionally, reports financial metrics, and supervises accounting staff. Requires extensive revenue accounting and billing experience in SaaS or software, management expertise, and strong technical accounting knowledge.
Top Skills: Ai-Enabled Accounting Automation ToolsBilling PlatformsErp SystemsExcelMicrosoft OutlookMicrosoft WordNetSuiteOracleSAP
2 Hours Ago
Remote or Hybrid
MN, USA
Senior level
Senior level
Software • Analytics • Hospitality
Leads revenue recognition, billing, accounts receivable, collections, audit readiness, and financial controls. Owns ASC 606 compliance, billing systems, process automation, AI-enabled accounting improvements, reporting, and cross-functional alignment. Manages and mentors the accounting team while supporting scalable revenue operations and public-company readiness.
Top Skills: Ai-Enabled Automation ToolsAsc 606GaapExcelMicrosoft OutlookMicrosoft WordNetSuiteOracleSAP

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account