Alljoined, Inc. Logo

Alljoined, Inc.

Data Infrastructure Engineer

Reposted 22 Days Ago
In-Office
San Francisco, CA, USA
140K-180K Annually
Mid level
In-Office
San Francisco, CA, USA
140K-180K Annually
Mid level
The Data Infrastructure Engineer will build backend architecture, manage data pipelines, and ensure high-performance compute clusters to support machine learning workloads.
The summary above was generated by AI
About Alljoined

Alljoined is creating a future where humans are fully understood and augmented by technology. Our work solves the communication bottleneck between humans and computers by decoding thoughts from the brain, entirely non-invasively. We apply deep learning research to large scale EEG datasets to decode multimedia input, eventually moving to internal thought. We are state-of-the art in capabilities and are fully vertically integrated. Our goal is to develop a general consumer interface to completely transform how we can live our lives.

We are actively growing our founding engineering team to build the underlying infrastructure that makes this ambitious future a reality.

About the Role

As a Data Infrastructure Engineer, you will build the backend and hardware architecture that allows us to do high-quality and fast research. You'll be owning our entire data lifecycle, from building pipelines that process massive multimodal datasets (video, audio, text, time-series) to provisioning and managing both cloud and bare metal compute clusters we use to train on it. You will be powering our foundational model training by bridging the gap between physical neuro hardware and our central repositories, working alongside world-class researchers to ensure they have a high-throughput, low-latency pipeline straight to the GPUs.

You might be a good fit if you
  • Have 3+ years of production software engineering experience with deep expertise in systems-level architecture and languages like Python, Rust, C++, or Go.

  • Have built and maintained high-performance ETL pipelines capable of processing, buffering, and storing terabytes of daily unstructured data.

  • Are comfortable architecting, provisioning, and maintaining bare-metal local compute clusters, storage servers, and high-speed networking for intensive ML workloads.

  • Have a background in handling continuous, highly concurrent data streams from heterogeneous hardware peripherals without data loss.

  • Are capable of working across hybrid environments to define storage topologies, manage databases (TimescaleDB, ClickHouse), and sync massive datasets between on-premise edge servers and the cloud (AWS/GCP/Azure).

  • Enjoy owning the entire technical lifecycle of infrastructure, from optimizing low-level I/O bound operations to production deployment.

Strong candidates may have
  • A deep understanding of modern ML frameworks (PyTorch/TensorFlow) and know how to build datasets that maximize and saturate GPU utilization.

  • Experience managing networking for distributed GPU training (InfiniBand, RoCE) or optimizing zero-copy networking and shared memory.

  • Built infrastructure involving programmatic video processing (FFmpeg, GStreamer, OpenCV)

Compensation Range

$140,000 - $180,000/year

While this represents our expected range based on market data, final compensation will be determined based on your specific skills and experience and may be outside this range.

Benefits
  • Competitive equity compensation at a seed stage startup

  • Options for housing support

  • Visa sponsorship

  • 3% 401k matching

  • Health insurance

HQ

Alljoined, Inc. San Francisco, California, USA Office

San Francisco, CA, United States

Similar Jobs

13 Days Ago
Remote or Hybrid
USA
100K-155K Annually
Senior level
100K-155K Annually
Senior level
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Build and operate AWS GovCloud data infrastructure across development, preproduction, and production environments. Establish PostgreSQL platforms, infrastructure-as-code with Chef, BI infrastructure, identity integrations, monitoring, observability, CI/CD, and disaster recovery. Support data pipelines, troubleshoot incidents, optimize databases, and ensure FedRAMP and FISMA compliance. Collaborate with security, compliance, DevOps, and data engineering teams while delivering greenfield infrastructure from architecture through production.
Top Skills: AWSAws GovcloudBashChefCi/CdCloudwatchDatadogEltETLGitGitlabGoogle SamlJenkinsNagiosPostgresPythonRubySsoTableau ServerVpc
2 Days Ago
In-Office or Remote
United States
285K-340K Annually
Entry level
285K-340K Annually
Entry level
Artificial Intelligence • Machine Learning • Natural Language Processing • Software • Generative AI
Design, build, and operate a distributed storage and data movement platform supporting petabyte-scale model training and evaluation. Run stateful systems across Kubernetes clusters, define throughput, latency, and durability requirements with research and infrastructure teams, and solve networking, I/O, consistency, and cross-region data transfer challenges. The role involves supporting cloud object storage, POSIX filesystems, model checkpoints, and large-scale GPU workloads.
Top Skills: Amazon S3Csi DriversGoKubernetesLustrePersistent VolumesPosix FilesystemsPythonStatefulsetsVastWeka
9 Days Ago
In-Office
San Francisco, CA, USA
500K-850K Annually
Senior level
500K-850K Annually
Senior level
Artificial Intelligence • Natural Language Processing • Generative AI
Design and operate scalable, reproducible data infrastructure for large language model pre-training. Build high-throughput distributed processing systems, including tokenization, deduplication, chunking, quality assurance, validation, and end-to-end pipelines that convert web-scale corpora into training-ready datasets. Collaborate with research teams on novel architectures while emphasizing reliability, fault tolerance, traceability, and performance.
Top Skills: SparkPythonRust

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account