Sciforium Logo

Sciforium

Pre-training Research Engineer

Posted 18 Days Ago
Be an Early Applicant
In-Office
San Francisco, CA, USA
165K-225K Annually
Senior level
In-Office
San Francisco, CA, USA
165K-225K Annually
Senior level
Develop and pre-train large byte-native and multimodal foundation models. Implement architectures, training objectives, optimization methods, and stable scaling recipes. Build production-grade training infrastructure, run ablations, analyze training dynamics, and improve model quality, efficiency, reliability, and scalability across GPU-based distributed environments.
The summary above was generated by AI

Sciforium is an AI infrastructure company developing next-generation multimodal AI models and a proprietary, high-efficiency serving platform. Backed by multi-million-dollar funding and direct sponsorship from AMD with hands-on support from AMD engineers the team is scaling rapidly to build the full stack powering frontier AI models and real-time applications.

About the role

As a Pre-training Research Engineer, you’ll focus on model implementation, pertaining and scaling, and improving the quality of our byte-native and multimodal foundation models. You’ll build and iterate quickly on research ideas, contribute production-grade training code and infrastructure, and help deliver high-quality base models that can serve real-world use cases at scale.

Key ResponsibilitiesPre-training & Scaling
  • Train large byte-native and multimodal foundation models across massive, heterogeneous corpora.

  • Implement and evaluate new model architectures, training objectives, and optimization methods.

  • Develop stable pre-training recipes and run scaling experiments for novel architectures.

  • Conduct ablations and analyze training dynamics, model behavior, and base-model quality.

  • Work with data and distributed training engineers to improve training efficiency, reliability, and scalability.

Must-Haves
  • 5+ years of experience in machine learning research or engineering, with a proven track record of developing and pre-training large language or multimodal foundation models.

  • Software Engineering: Strong general software engineering skills, with the ability to write robust and performant training code.

  • ML Foundations: Solid understanding of deep learning fundamentals and modern pre-training methods and literature.

  • Research and Experimentation: Ability to quickly implement research ideas and evaluate them using clear baselines, ablations, metrics, and analysis.

  • GPU and Distributed Training: Hands-on experience running training workloads in GPU-based environments, with familiarity with distributed training.

  • Education: MS in Computer Science, Machine Learning, Artificial Intelligence, Mathematics, or a related field.

Nice-to-Haves
  • PhD in Computer Science, Machine Learning, Artificial Intelligence, Mathematics, or a related field.

  • JAX Ecosystem: Extensive experience with the JAX, Flax, and XLA stack.

  • Large-Scale Distributed Training: Experience with multi-node pre-training using systems such as FSDP, ZeRO, or Megatron.

  • Training Recipes and Scaling: Experience developing training recipes, ablations, or scaling experiments.

  • Monitoring and Reproducibility: Experience owning end-to-end training and evaluation pipelines with monitoring and reproducibility.

Education
  • MS or PhD in Computer Science, Machine Learning, Artificial Intelligence, Mathematics, or a related field.

Benefits include
  • Medical, dental, and vision insurance

  • 401k plan

  • Daily lunch, snacks, and beverages

  • Flexible time off

  • Competitive salary and equity

Equal opportunity

Sciforium is an equal opportunity employer. All applicants will be considered for employment without attention to race, color, religion, sex, sexual orientation, gender identity, national origin, veteran or disability status.

HQ

Sciforium San Francisco, California, USA Office

San Francisco, CA, United States

Sciforium Los Altos, California, USA Office

4401 El Camino Real, Los Altos, California, United States, 94022

Similar Jobs

18 Days Ago
In-Office or Remote
San Francisco, CA, USA
Entry level
Entry level
Artificial Intelligence • Information Technology • Software
Build and optimize decentralized distributed pretraining systems for frontier-scale models on heterogeneous consumer hardware. Responsibilities include model-parallel training, communication efficiency, elasticity, fault tolerance, checkpointing, state synchronization, recovery, and monitoring across unreliable internet-connected devices. The role requires hands-on PyTorch distributed training experience, strong production Python skills, concurrency and profiling expertise, and evidence of shipping research or engineering systems.
Top Skills: Data ParallelismDeepspeedDistributed TrainingFsdpMegatronNat TraversalP2P NetworkingPipeline ParallelismPythonPyTorchTensor Parallelism
27 Days Ago
In-Office or Remote
San Francisco, CA, USA
350K-850K Annually
Mid level
350K-850K Annually
Mid level
Artificial Intelligence • Natural Language Processing • Generative AI
The Research Engineer will develop large language models, conduct research on model architecture, optimize training infrastructure, and lead research projects.
Top Skills: Deep Learning FrameworksEtl ProcessesKubernetesLarge-Scale Machine LearningPythonPyTorch
17 Minutes Ago
Hybrid
San Francisco, CA, USA
169K-361K Annually
Senior level
169K-361K Annually
Senior level
Consumer Web • Coupons • Healthtech • Social Impact • Pharmaceutical
Leads and scales AI engineering teams responsible for production LLM systems, agents, integrations, and shared AI platforms. Oversees architecture, delivery, reliability, cost, latency, evaluation, observability, governance, security, and operational practices. Partners with product and business leaders to prioritize investments, manage dependencies, and establish reusable engineering standards. Develops engineering leaders and drives alignment across teams while ensuring safe, scalable AI adoption.
Top Skills: Ai AgentsAPIsCloud InfrastructureFine-TuningLlmLlmopsMcpMlopsNist Ai RmfPrompt EngineeringRag

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account