Metasys Logo

Metasys

Cloud Architect Internship

Posted Yesterday
Be an Early Applicant
Remote
Hiring Remotely in United States
Internship
Remote
Hiring Remotely in United States
Internship
Design and govern a multi-tenant cloud infrastructure (AWS/GCP/Azure) for an e-commerce platform. Optimize capacity and costs for fluctuating traffic and AI agent compute, implement multi-region HA/DR for PostgreSQL 15 and Redis, enforce Terraform IaC governance, ensure cloud security and data residency compliance, and integrate containerized apps (Docker/Traefik) with monitoring (Prometheus, OpenTelemetry).
The summary above was generated by AI
Overview: Multi-Tenant Cloud Infrastructure Design

The Cloud Architect is responsible for designing, managing, and optimizing the entire multi-tenant cloud infrastructure that hosts our integrated supply chain e-commerce platform. While our current stack utilizes Oracle Cloud Free VMs, this role requires expertise in major public clouds (AWS/GCP/Azure) to architect scalable, resilient, cost-optimized, and compliant solutions that handle fluctuating e-commerce traffic, stable internal tool usage, and intensive AI agent compute demands.

Internship Details

Duration: 3 months
Start Date: Immediate
Location: Remote
Stipend: None initially. Based on your first-quarter performance, you may be offered a paid full-time opportunity, or even be absorbed directly by the client as an FTE.

Key Responsibilities & Core Projects

You will design and govern the platform's foundation, ensuring scalability and compliance across all environments.

  • Cloud Architecture Design: Design and evolve the target cloud infrastructure (utilizing AWS, GCP, or Azure best practices) for maximum scalability, security, and high-availability, ensuring the platform can reliably support the MES $\rightarrow$ WMS $\rightarrow$ OMS flow.

  • Capacity Planning & Optimization: Plan and optimize resource allocation to effectively handle unpredictable e-commerce traffic spikes and the specific compute requirements for AI agent training and inference, driving cost efficiency.

  • High Availability & Disaster Recovery (DR): Implement multi-region/multi-AZ high-availability architectures and define comprehensive Disaster Recovery (DR) strategies for all core services, including PostgreSQL 15 and Redis.

  • Infrastructure-as-Code (IaC) Governance: Establish and enforce best practices for Terraform usage, ensuring configuration consistency, security compliance, and auditable infrastructure changes.

  • Security & Compliance: Conduct regular security reviews of cloud configurations (e.g., IAM, VPC/VNet, Storage) and ensure architecture aligns with data residency and compliance requirements relevant to our multi-tenant operations.

  • Service Integration: Architect the network and service mesh overlay that integrates the containerized applications (Docker/Traefik) with external cloud services and the overall monitoring solution (Prometheus, OpenTelemetry).

Required Technologies & Tools

Candidates must possess deep architectural experience with public cloud providers and infrastructure automation:

  • Cloud Providers: Expert-level proficiency in at least one major public cloud (AWS, GCP, or Azure).

  • Infrastructure-as-Code (IaC): Mandatory expertise in Terraform for cloud resource provisioning.

  • Containerization: Deep knowledge of Docker networking, security, and orchestration principles.

  • Networking & Edge: Experience configuring load balancing, service mesh, and ingress controllers (e.g., Traefik).

  • Data & Storage: Architecting scalable database services (PostgreSQL) and object storage (MinIO).

  • Security: Cloud security best practices, IAM policy design, and network segmentation.

AI Agent Focus

You will optimize the cloud layer for our emerging AI capabilities.

  • Compute Optimization: Design elastic and cost-effective compute clusters (e.g., GPU instances) to efficiently handle the variable demands of LLM fine-tuning and multi-agent system orchestration.

  • Data Residency: Architect the data pipeline and storage solutions to ensure that training data and model artifacts adhere to strict data residency requirements across tenants.

Success Metrics & Career Path

Performance will be measured by:

  • Cost Efficiency: Demonstrable reduction in cloud operational costs (FinOps) while maintaining performance.

  • Availability: Achieving defined SLAs/SLOs for infrastructure uptime and performance.

  • Compliance: Successful implementation and auditing of cloud security and data residency controls.

Mentorship Structure: Reports to the Solution Architect or Head of Technology, collaborating closely with the SRE and DevSecOps teams to operationalize cloud strategy.

Similar Jobs

10 Minutes Ago
Remote or Hybrid
111K-145K Annually
Senior level
111K-145K Annually
Senior level
Consumer Web • eCommerce • Machine Learning • Software • Sports • Analytics
The Senior Real Estate Project Manager in APAC will lead capital buildouts and renovations, manage project budgets, and collaborate with various stakeholders while ensuring design and operational consistency.
15 Minutes Ago
Remote
USA
Senior level
Senior level
Aerospace • Hardware • Software • Virtual Reality • Defense
Lead C2-focused business development for military training markets: build and manage pipeline, shape requirements, capture funded programs, establish senior military relationships, and align cross-functional teams to transition AR/LVC training technologies into scalable DoD programs.
Top Skills: ArAtarsDistributed Mission TrainingIsrLvcModeling And SimulationXr
26 Minutes Ago
Easy Apply
Remote
United States
Easy Apply
204K-290K Annually
Senior level
204K-290K Annually
Senior level
Big Data • Fintech • Mobile • Payments • Financial Services
Drive technical strategy and roadmap for Affirm's Online Storage platform. Design and build multi-region, highly available datastore solutions and control planes for hundreds of databases. Automate schema migrations, disaster recovery, sharding, and performance tuning. Collaborate with product, SRE, and infrastructure teams, own operations and on-call readiness, and mentor engineers while setting code and design standards.
Top Skills: AWSDistributed SqlDynamoDBKotlinKubernetesMySQLPgbouncerPostgresProxysqlPythonRds ProxyRedisSparkTidbVitess

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account