Krea Logo

Krea

Engineer, Supercomputing & Distributed Systems

Reposted One Month Ago
Be an Early Applicant
In-Office
San Francisco, CA, USA
Entry level
In-Office
San Francisco, CA, USA
Entry level
The role involves designing and managing distributed systems infrastructure for AI workloads, optimizing data pipelines, and collaborating on machine learning projects.
The summary above was generated by AI
About Krea

At Krea, we are building next-generation AI creative tools.

We're dedicated to making AI intuitive and controllable for creatives - our mission is to build tools that empower human creativity, not replace it. We believe AI is a new medium that allows us to express ourselves through various formats - text, images, video, sound, and even 3D. We're building better, smarter, and more controllable tools to harness this medium. We recently took this a step forward with the launch of Krea 2, our first foundation model, built completely from scratch for aesthetic diversity and stylistic control.

We've raised over $83M and are backed by world-class investors such as a16z, Bain Capital, and Abstract. We work full-time and in-person at our waterfront office in San Francisco. We care about creativity: our team includes musicians, designers, visual artists, and engineers.

 
Supercomputing / AI Infra at Krea

We build and operate the infrastructure for Krea's research and inference. Distributed training, 1000+ K8s GPU clusters, petabyte scale data pipelines, etc. We build a lot of this from scratch — custom distributed datastores, job orchestration systems, and streaming pipelines that replace tools like Kafka and Ray for modern AI workloads at scale.

Example projects:Distributed data systems
  • Design multi-stage pipelines that turn petabytes of raw data into clean, annotated datasets

  • Run classification models on billions of images

  • Deploy and combine LLMs to caption massive multimedia data

GPU infrastructure
  • Manage distributed training and inference on 1000+ GPU Kubernetes clusters

  • Solve orchestration and scaling for large-scale GPU job processing

  • Scale workloads and research between clusters in multiple datacenters

Distributed training
  • Profile and optimize dataloaders streaming thousands of images per second

  • Profile and debug InfiniBand networking on huge training runs

  • Build fault tolerance systems for large-scale pretraining

  • Collaborate with researchers on evolving RL infrastructure

Applied ML pipelines
  • Find clean scenes in millions of videos using distributed shot-boundary detection

  • Customize and train models to filter billions of images for questions like "is this a screenshot?"

  • Build the systems that bridge raw cluster capacity and research output

Who we're looking for

Systems people. If you've read a blog post about InfiniBand debugging or building a custom distributed database and thought "I want to do that" — this is that team.

You'll spend your time working heavily with Python, Kubernetes, Torch, and data tools like DuckDB, Arrow, etc. It's OK if you don't have K8s or ML experience — the main thing we hire for is an intuition for distributed systems, and a great mental model of how systems interact and function under different conditions.

Strong candidates may have experience with…
  • Python, PyArrow, DuckDB, SQL, massive relational databases, PyTorch, Pandas, NumPy…

  • Kubernetes

  • Designing and implementing large-scale ETL systems

  • Fundamental knowledge of containerization, operating systems, file-systems, and networking

  • Distributed systems design

  • Distributed training systems (NCCL, InfiniBand, RDMA)

  • Streaming and event processing systems (Kafka, Pulsar, or similar)

  • PyTorch internals, custom dataloaders, and training infrastructure

What we offer
  • Team: Work alongside a world-class team building the future of AI creative tooling

  • Impact: Significant scope and company-wide impact

  • Competitive compensation: generous salary & equity packages

  • Health & wellness: 100% health & 99% dental/vision insurance premiums covered for employees, health FSA accounts, & long-term disability coverage

  • Time off: Flexible PTO policy

  • Financial planning: 401k with a 4% company-sponsored match

  • Meals in the office: breakfast, lunch, dinner - you name it, we'll cover it

  • Transit: Ubers covered to & from the office

  • Sponsorship: We're open to sponsoring international visas where we can (e.g., STEM OPT, OPT, H-1B, O-1, E-3).

  • And more!

Please note the above benefits & perks are for full-time employees

Similar Jobs

58 Seconds Ago
Remote or Hybrid
United States
98K-164K Annually
Expert/Leader
98K-164K Annually
Expert/Leader
Artificial Intelligence • Cloud • Sales • Security • Software • Cybersecurity • Data Privacy
Leads a team of Forward Deployed Engineers while supporting sales, customer discovery, solution architecture, foundational implementations, and proof-of-value engagements. The role translates customer requirements into SailPoint solutions, advises on product architecture and best practices, collaborates across sales, product, services, and customer success, and develops AI-powered agentic solutions for demonstrations and customer outcomes. It requires significant enterprise software implementation or presales experience, technical leadership, SaaS architecture knowledge, and client-facing consulting expertise.
Top Skills: Active DirectoryAgentic AiAngularArtificial IntelligenceAWSAzureCassandraClaude CodeCursorGCPGithub CopilotGoogle AntigravityJ2EeJavaJavaScriptJSONKiroLdapLinuxMicrosoft Sql ServerMongoDBMySQLNode.jsOne IdentityOpenai CodexOracleOracle Identity ManagerPeoplesoftPowershellPythonReactRedisRsa AveksaSailpointSAPSaviyntServicenowSQLSybaseTypescriptUnixWindowsXML
3 Minutes Ago
Remote or Hybrid
San Francisco, CA, USA
127K-271K Annually
Senior level
127K-271K Annually
Senior level
Consumer Web • Coupons • Healthtech • Social Impact • Pharmaceutical
Leads integrated B2B marketing strategy across healthcare professional, pharmaceutical, pharmacy, employer, payer, and health technology segments. Owns HCP lifecycle marketing, messaging, content, campaigns, commercialization, cross-functional programs, measurement, and executive recommendations. Develops collateral and governance, coordinates stakeholders across marketing, sales, analytics, communications, and commercial teams, and supports industry events. Drives scalable marketing practices, audience engagement, revenue growth, and measurable ROI.
5 Minutes Ago
Easy Apply
Hybrid
San Francisco, CA, USA
Easy Apply
105K-234K Annually
Senior level
105K-234K Annually
Senior level
Fintech • Information Technology • Payments • Productivity • Software • Travel • Automation
Lead marketing analytics: build dashboards, attribution and predictive models, ensure data governance and quality, drive self-service reporting, conduct deep-dive campaign and funnel analyses, and partner with marketing and finance to inform strategy and operations.
Top Skills: AmplitudeApi ManagementDbtGa4GitGoogle Apps ScriptGoogle SheetsMarketoSalesforceSQLThoughtspot

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account