Engramme Logo

Engramme

ML Systems & Performance Engineer

Posted Yesterday
In-Office
San Francisco, CA, USA
Senior level
In-Office
San Francisco, CA, USA
Senior level
Design, optimize, and scale training and inference workloads for personalization and continual learning. Productionize research prototypes, improve latency/reliability/efficiency, optimize per-user serving and memory retrieval, and work on distributed training across GPUs and nodes. Collaborate with researchers, platform engineers, and customers; establish engineering practices, testing, on-call, and security posture.
The summary above was generated by AI
About Engram

Today’s AI is a brilliant stranger: it can solve the world’s hardest math problems, but it knows next to nothing about you and your work. It rereads your files to answer even basic questions, burns an enormous amount of tokens when sifting through large corpuses, and between sessions, it retains scraps at best.

We train models to study your world and anticipate your questions in advance, forming engrams: compact memories that capture your knowledge and history. Our approach opens a new axis of scaling. The more we study your context at training time, the better we become at inference time.

We're already working with leaders in AI like Microsoft, Notion, and Harvey, and just raised $98M from General Catalyst, Kleiner Perkins, Sequoia, Factory, Modern, Amplify, Neo and others. Our investors and advisors include Assaf Rappaport, Andrej Karpathy, and Pieter Abbeel.

AI has spent years learning everything about the world. Now it should learn something about yours.

About this role

You’ll be among the first ML Systems Engineer, joining a team of machine learning researchers and performance engineers in building our personalization and continual learning API, which powers models and agents that learn from user context.

This role is focused on designing, optimizing, and scaling training and inference workloads —bridging the gap between cutting-edge AI research and production. This includes:

  • Designing and executing new frameworks, techniques, and systems to improve performance, reliability, latency, and efficiency.

  • Partner closely with researchers, turning prototypes into systems that run at scale and feeding systems constraints back into research decisions.

  • Optimize serving paths for personalization and memory retrieval, where per-user state and low latency both matter.

  • Work on distributed training — data and model parallelism, communication scheduling, and scaling efficiency across multiple GPUs and nodes.

You'll be the bridge between researchers and platform engineering, while working in deep collaboration with our customers (AI-native application-layer companies like Notion and Harvey). The architecture will be shaped by the constraints and requirements of our R&D work. There are no walls between product, research, and engineering here; delivering on our mission requires a multidisciplinary approach.

This is a founding hire in the truest sense. You’ll set the bar for engineering at Engram: code review, testing, on-call, and a security posture that gives customers confidence in entrusting us with their most sensitive data. You’ll also help build the engineering team around you and influence our engineering culture as we scale.

Your background looks like
  • Bachelor’s degree or equivalent experience in computer science, engineering, or similar.

  • 5+ years of experience with training or inference systems, optimized workloads with measurable results.

  • Strong engineering foundation, with demonstrated excellence navigating complex technical environments and shipping high-quality code in a fast-paced environment.

  • Deep understanding of ML framework (eg. PyTorch, JAX), GPUs, distributed systems, and infrastructure.

  • Operate well in ambiguous environments — you will have real ownership and be responsible for steering the ship in a novel sector of the industry.

  • You have a bias toward action and a knack for turning research concepts into concrete, executable plans.

Bonus points if you have
  • Prior early-stage experience.

  • Experience in open-source ML or systems infrastructure projects.

Engram is based in San Francisco. This role is in-person in our SF office. We offer competitive cash compensation and startup equity.

Engram is an equal opportunity employer. We’re building a team that reflects a range of backgrounds and perspectives, and we welcome applicants regardless of race, color, religion, national origin, gender, gender identity, sexual orientation, age, disability, or veteran status.

Similar Jobs

22 Days Ago
In-Office
Sunnyvale, CA, USA
Mid level
Mid level
Artificial Intelligence
The Performance Engineer - Inference will optimize model inference speed and throughput, debug low-level kernel performance, and develop tools to visualize performance data.
Top Skills: C++Python
15 Days Ago
In-Office
San Francisco, CA, USA
180K-250K Annually
Mid level
180K-250K Annually
Mid level
Cloud • Digital Media • Information Technology
Design and implement model serving architectures, develop monitoring tools, and optimize performance for generative media models working with Applied ML teams.
Top Skills: NsightPyTorchTensorrtTransformerengineTriton
5 Minutes Ago
Hybrid
124K-335K Annually
Senior level
124K-335K Annually
Senior level
Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI
Lead tax operations transformation engagements: advise clients on tax reporting strategies and compliance, analyze financial data, coach teams, apply systems thinking, and deliver clear client-facing recommendations while upholding professional standards.

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account