Cox Exponential Jobs

Founding Engineer, AI Infra

Cox Exponential

Founding Engineer, AI Infra

Reposted 20 Days Ago

Remote or Hybrid

7 Locations

Senior level

Remote or Hybrid

7 Locations

Senior level

Design, build, and operate end-to-end training and inference infrastructure for large language and multimodal models. Improve efficiency (memory, parallelism, kernel optimizations), ensure robust scalable training and RL pipelines, optimize low-latency/high-throughput serving (quantization, caching, speculative decoding), manage multi-GPU and multi-cloud orchestration, and productionize new algorithms with strong observability and reproducibility.

The summary above was generated by AI

About Goaly

At Goaly, our mission is to make custom AI affordable for every business. Our founding team comes from the front lines of top AI labs and tech giants (Meta MSL, TikTok AI, Google DeepMind, xAI, Microsoft Research, etc.), where we built large-scale training infrastructure powering trillion-parameter models and scaled GenAI models to a global user base. Now, we are building something we wish we had before: a platform that makes training and adapting custom AI affordable for all modern companies, not just Big Tech. Our north star is ambitious: for a domain-specific task, reach 90% of SOTA performance at less than 10% of the cost. To get a taste of what we are doing, see our first tech blog.

About the Role

You will sit at the intersection of systems engineering and applied ML, building specialized infrastructure that keeps large language and multimodal models fast, reliable, and cost-effective. You will partner with research, product, and infra teams to ship production-ready platforms for training and serving AI at scale.

Key Responsibilities

Efficiency & performance: Improve LLM training and inference efficiency through better memory utilization, optimized parallelism, and kernel-level innovations (e.g. FlashAttention, CUDA/Triton).
Training & RL robustness: Build scalable, stable training and RL pipelines with strong reproducibility, observability, and debuggability.
Serving & inference optimization: Design and tune high-throughput, low-latency model serving systems, including quantization, caching, and speculative decoding.
Scalability & infrastructure: Own end-to-end training and inference infrastructure — from data ingestion and checkpointing to multi-GPU and multi-cloud orchestration.
Production enablement: Work closely with researchers and product engineers to turn new algorithms into reliable, production-ready systems.

Requirements

5+ years building or operating ML infrastructure at scale, ideally supporting large language or multimodal models.
Deep understanding of GPU architecture, distributed training frameworks (PyTorch, DeepSpeed, Megatron, Ray), and parallelism strategies.
Hands-on experience running inference stacks (vLLM / SGLang, TGI, Triton) and optimizing them via low-level profiling.
Strong software engineering fundamentals in Python and one of C++/Rust/Go, with clean, reliable code shipped to production.
Working knowledge of modern data pipelines, feature stores, and vector databases used in production AI systems.
Comfort automating infrastructure with Kubernetes, Terraform/Pulumi, and observability stacks (Prometheus, Grafana, OpenTelemetry).

Bonus Points

Experience deploying open-source LLMs (Llama 3, Qwen, DeepSeek) or training custom foundation models.
Contributions to ML systems tooling (compilers, kernels, inference runtimes) or open-source infrastructure projects.
Background in reinforcement learning, evaluation harnesses, or alignment tooling that hardens production AI systems.

Similar Jobs

Square

Account Executive

3 Hours Ago

Remote or Hybrid

149K-223K Annually

Mid level

149K-223K Annually

Mid level

eCommerce • Fintech • Hardware • Payments • Software • Financial Services

Field-driven territory sales role responsible for full-cycle, self-sourced selling: prospecting, delivering live demos, closing deals, and building pipeline through door-to-door outreach, partnerships, and events. Spend most weeks in-market, manage Salesforce activity, master verticals (restaurants, retail, services), and consistently exceed quota while ensuring smooth onboarding and strong local presence.

Top Skills: Salesforce

Block

Account Executive

13 Hours Ago

In-Office or Remote

149K-223K Annually

Mid level

149K-223K Annually

Mid level

Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency

Field-driven, full-cycle sales role responsible for building pipeline, conducting in-person demos, closing deals, and establishing Square as the go-to solution in the territory. Drive 50–60 targeted weekly visits, develop partnerships, manage Salesforce pipeline, and exceed quota while specializing in restaurants, retail, and services verticals.

Top Skills: SalesforceSquare

DuckDuckGo

Senior Back-end Engineer

13 Hours Ago

Remote

USA

179K-179K Annually

Senior level

179K-179K Annually

Senior level

Information Technology

As a Senior Backend Engineer at DuckDuckGo, you'll lead backend projects, mentor engineers, and develop AI-enhanced features for the company's privacy-centric product line.

Top Skills: Ai ToolingGoNode.jsPerlRag Pipelines

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
Major Tech Employers: Google, Apple, Salesforce, Meta
Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine