Lambda Logo

Lambda

Senior Software Engineer - Infrastructure Storage

Posted 10 Hours Ago
Remote or Hybrid
2 Locations
266K-395K Annually
Senior level
Remote or Hybrid
2 Locations
266K-395K Annually
Senior level
Design, develop, and maintain high-performance distributed storage software and protocols (file, block, object). Integrate with hardware (NVMe, DPUs, GPU-direct), optimize performance and scalability, troubleshoot production data centers, collaborate with networking, compute, control plane and observability teams, and lead/mentor engineering teams.
The summary above was generated by AI

Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence. One person, one GPU.

If you'd like to build the world's best AI cloud, join us.

*Note: This position requires presence in our San Francisco office location 4 days per week; Lambda’s designated work from home day is currently Tuesday.

In the world of distributed AI, raw GPU and CPU horsepower is just a part of the story. High-performance networking and storage are the critical components that enable and unite these systems, making groundbreaking AI training and inference possible.

The Lambda Infrastructure Engineering organization forges the foundation of high-performance AI clusters by welding together the latest in AI storage, networking, GPU and CPU hardware.

Our expertise lies at the intersection of:

  • High-Performance Distributed Storage Solutions and Protocols: We engineer the protocols and systems that serve massive datasets at the speeds demanded by modern clustered GPUs.

  • Dynamic Networking: We design advanced networks that provide multi-tenant security and intelligent routing without compromising performance, using the latest in AI networking hardware.

  • Compute Virtualization: We enable cutting-edge virtualization and clustering that allows AI researchers and engineers to focus on AI workloads, not AI infrastructure, unleashing the full compute bandwidth of clustered GPUs.

About the Role:

We are seeking a seasoned Storage Software Engineer with experience designing and deploying various storage protocol solutions at scale (object, block, and file).

This is a unique opportunity to work at the intersection of large-scale distributed systems and the rapidly evolving field of artificial intelligence infrastructure. This is an opportunity to have a significant impact on the future of AI. You will be building the foundational infrastructure that powers some of the most advanced AI research and products in the world.

What You’ll Do

  • Technical Leadership:

    • At the Senior Level

  • Execution:

    • Systems-Level Programming and Architecture

    • Design, develop, and maintain software for storage systems, focusing on performance, scalability, and reliability.

    • Implement and optimize storage protocol APIs for file (e.g., NFS, SMB), block (e.g., Fibre Channel), and object (e.g., S3) access.

    • Develop distributed systems for managing and orchestrating storage resources across multiple storage solutions and redundant arrays.

    • Collaborate with hardware and system architects to integrate software with various storage solutions, including NVMe and GPU-direct storage.

    • Troubleshoot and debug complex issues in a production data center environment.

    • Contribute to the full software development lifecycle, from requirements gathering and design to deployment and maintenance.

  • Collaboration

    • Work closely with the storage software teams and networking teams to execute on cross-functional infrastructure initiatives and new data-center deployments including integration of storage protocols across a variety of on-prem storage solutions.

    • Work closely with the control plane and MK8s teams to meet customer/product requirements for usability, reliability, and telemetry.

    • Work with the observability team to build/track SLOs/SLIs.

    • Work closely with Networking, Compute, and Storage Software Engineering teams to deploy high-performance distributed storage solutions to serve AI/ML workloads.

    • Partner with the fleet engineering team to ensure seamless deployment, monitoring, and maintenance of the distributed storage solutions.

  • Innovate:

    • Stay current with the latest trends and research into AI and HPC storage technologies.

    • Work with the Lambda product team to uncover new trends in the AI inference and training product category that will inform emerging storage solutions.

    • Optimize protocol solutions for the AI product vertical exploring optimizations for AI Inference, training, and scientific computing applications.

You

  • Experience:

    • 10+ years of experience in storage engineering with at least 5+ years in a management or lead role.

    • Systems-Level Programming and Architecture

    • Storage Protocol and API Mastery:

    • Storage Performance Optimization

    • DPKD SPKD

    • Physical Infrastructure Knowledge

    • Operational Acumen

  • Technical Skills:

    • Experience in serving one or more of the following storage protocols: object storage (e.g., S3), block storage (e.g., iSCSI), or file storage (e.g., NFS, SMB, Lustre).

    • Professional individual contributor experience as a storage engineer or storage SRE.

    • Familiarity with modern storage technologies (e.g., NVMe, RDMA, DPUs) and their role in optimizing performance.

  • People Management:

    • Experience building a high-performance team through deliberate hiring, upskilling, planned skills redundancy, performance-management, and expectation setting.

Nice to Have

  • Experience:

    • Experience driving cross-functional engineering management initiatives (coordinating events, strategic planning, coordinating large projects).

    • Experience with NVidia SuperNIC DPUs for edge-caching (such as implementing GPUDirect Storage).

  • Technical Skills:

    • Deep experience with Vast, Weka and/or NetApp in an HPC or AI Infrastructure environment.

    • Deep experience implementing CEPH in an HPC or AI infrastructure environment at a scale greater than 100PB.

  • People Management:

    • Experience driving organizational improvements (processes, systems, etc.)

    • Experience training, or managing managers.

Salary Range Information

The annual salary range for this position has been set based on market data and other factors. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.

About Lambda

  • Founded in 2012, with 500+ employees, and growing fast

  • Our investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent Cove

  • We have research papers accepted at top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOG

  • Our values are publicly available: https://lambda.ai/careers

  • We offer generous cash & equity compensation

  • Health, dental, and vision coverage for you and your dependents

  • Wellness and commuter stipends for select roles

  • 401k Plan with 2% company match (USA employees)

  • Flexible paid time off plan that we all actually use

A Final Note:

You do not need to match all of the listed expectations to apply for this position. We are committed to building a team with a variety of backgrounds, experiences, and skills.

Equal Opportunity Employer

Lambda is an Equal Opportunity employer. Applicants are considered without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.

HQ

Lambda San Francisco, California, USA Office

San Francisco, CA, United States, 94107

Similar Jobs

7 Days Ago
Remote
United States
203K-274K Annually
Senior level
203K-274K Annually
Senior level
Artificial Intelligence • Cloud • Consumer Web • Productivity • Software • App development • Data Privacy
Design, implement, and operate large-scale distributed storage systems ensuring data durability, availability, and performance. Build and maintain replication, erasure coding, and lifecycle systems; write high-quality Go/Rust/C++ code; participate in on-call rotations; troubleshoot production incidents; collaborate with networking, hardware, and capacity teams; lead scoped projects and drive infrastructure evolution.
Top Skills: C++CephColossusErasure CodingGfsGoReplication ProtocolsRustS3
3 Days Ago
Remote
United States
172K-215K Annually
Senior level
172K-215K Annually
Senior level
Aerospace • Big Data • Greentech • Hardware • Social Impact
Design, build, and operate Planet's scalable object and metadata storage platform. Own backend services, Elasticsearch indexing, search, event notification, and disaster recovery. Measure performance, create alerts, be on call, drive system design and architecture, mentor engineers, and contribute to the technical roadmap for highly available, cloud-native data infrastructure.
Top Skills: AWSBigtableClaudeCopilotElasticsearchGCPGoKubernetesMySQLPostgresPythonTerraform
13 Days Ago
Remote
United States of America
128K-267K Annually
Senior level
128K-267K Annually
Senior level
AdTech • Digital Media • Information Technology • Other
Design and optimize storage systems at scale, implementing caching strategies and ensuring data availability and performance for a large user base.
Top Skills: AWSCloud SpannerGCPGoJavaMemcachedPythonRedisSQLValkey

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account