Lightning AI

HQ
New York
Total Offices: 2
50 Total Employees
Year Founded: 2019

Jobs at Lightning AI

Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.

Recently posted jobs

4 Hours AgoSaved
Hybrid
3 Locations
Artificial Intelligence • Machine Learning • Software
Build and operate backend services, control planes, and automation for Lightning AI’s managed Kubernetes and Slurm infrastructure. Develop distributed systems for cluster provisioning, lifecycle management, scheduling, observability, reliability, and scalability across large-scale GPU environments. Diagnose complex production issues and collaborate with infrastructure, AI, and platform engineering teams. Contribute to architecture, technical design, mentoring, engineering practices, and on-call operations.
4 Hours AgoSaved
Hybrid
San Francisco, CA, USA
Artificial Intelligence • Machine Learning • Software
Own and scale Lightning AI’s global corporate tax function across a multi-entity organization. Lead federal, state, local, indirect, and property tax compliance; ASC 740 tax accounting; forecasting; audit support; M&A tax; risk management; and process automation. Build policies, controls, documentation, and a scalable tax operating roadmap while partnering with Finance, Accounting, FP&A, Legal, auditors, and external advisors.
4 Hours AgoSaved
Hybrid
San Francisco, CA, USA
Artificial Intelligence • Machine Learning • Software
Support multi-entity general ledger close activities, including journal entries, reconciliations, intercompany, payroll accruals, prepaids, fixed assets, leases, financial statements, and flux analysis. Help transform and document the close function through NetSuite implementation, process redesign, and control development. Support external audits and maintain accurate, auditable accounting records.
4 Hours AgoSaved
Remote or Hybrid
3 Locations
Artificial Intelligence • Machine Learning • Software
Designs, deploys, automates, and operates large-scale NVIDIA InfiniBand fabrics for GPU clusters. Manages UFM, switches, firmware, congestion control, routing, QoS, observability, troubleshooting, capacity planning, and incident response. Collaborates with AI platform, GPU infrastructure, storage, and systems teams to build reliable AI Factory environments.
4 Hours AgoSaved
Hybrid
3 Locations
Artificial Intelligence • Machine Learning • Software
Design, build, and maintain scalable backend services and APIs powering Lightning AI Studio. Develop platform capabilities for authentication, resource management, multi-tenancy, and developer workflows. Improve reliability, observability, scalability, and performance across cloud-native infrastructure. Collaborate with product, infrastructure, and AI engineering teams, lead technical initiatives from architecture through production, evaluate engineering processes, and mentor engineers.
4 Hours AgoSaved
Hybrid
3 Locations
Artificial Intelligence • Machine Learning • Software
Design and build backend systems for AI agent orchestration, tool execution, workflow management, memory, and state handling. Develop scalable APIs and reliable infrastructure for distributed AI workflows, improve platform reliability and observability, and partner with product, research, and engineering teams to deliver agent capabilities. Own technical initiatives from architecture through production, maintain software quality through testing and continuous delivery, and mentor engineers on backend and distributed-systems practices.
4 Hours AgoSaved
Hybrid
4 Locations
Artificial Intelligence • Machine Learning • Software
Develop and scale Lightning AI’s backend platform using Go, spanning APIs, infrastructure, billing, security, and integrations. Own features end to end, design reliable and scalable systems, improve architecture and performance, maintain software quality and continuous delivery, reduce technical debt, and mentor engineers. Collaborate with engineering, product, and design teams in a fast-changing SaaS environment.
4 Hours AgoSaved
Hybrid
4 Locations
Artificial Intelligence • Machine Learning • Software
Build and deploy production AI systems for customers, translating business objectives into scalable technical solutions. Responsibilities include architecture, proof-of-concepts, software development, deployment, monitoring, debugging, inference optimization, and distributed systems operation. The engineer partners with customer engineering teams, collaborates with product and engineering, improves reusable platform capabilities, and owns technical engagements from discovery through production scaling.
4 Hours AgoSaved
Remote or Hybrid
4 Locations
Artificial Intelligence • Machine Learning • Software
Build and operate production software, APIs, tooling, and automation for large-scale GPU, bare-metal, and HPC infrastructure. Responsibilities include provisioning, configuration, monitoring, lifecycle management, observability, hardware integration, reliability improvements, and infrastructure capacity deployment. The role partners with networking, data center, platform, and infrastructure teams to design scalable systems and define technical direction.
4 Hours AgoSaved
Hybrid
2 Locations
Artificial Intelligence • Machine Learning • Software
Designs, implements, and maintains secure network infrastructure for Lightning AI’s production and AI environments. Responsibilities include firewall, VPN, IDS/IPS, segmentation, vulnerability assessment, penetration testing, SIEM monitoring, incident response, compliance documentation, access controls, and security automation. The role partners with infrastructure and engineering teams to mitigate threats and optimize security controls.
4 Hours AgoSaved
Remote or Hybrid
4 Locations
Artificial Intelligence • Machine Learning • Software
This is a general talent community opportunity rather than a specific open position. Lightning AI invites candidates to submit their information for consideration as future roles become available across U.S. and London hubs, with occasional remote opportunities. The company develops tools and infrastructure for building, training, and deploying AI systems and values ownership, urgency, communication, teamwork, continuous improvement, and long-term thinking.
4 Hours AgoSaved
Remote or Hybrid
3 Locations
Artificial Intelligence • Machine Learning • Software
Design, optimize, and deploy large language model training and post-training pipelines. Improve model quality through fine-tuning, reinforcement learning, preference optimization, evaluation, and experimentation. Build PyTorch-based infrastructure, optimize distributed multi-GPU training, diagnose performance and convergence issues, and develop production-ready AI systems. Collaborate with researchers, infrastructure engineers, platform teams, and customers while contributing to open-source projects and reusable training capabilities.
4 Hours AgoSaved
Remote or Hybrid
4 Locations
Artificial Intelligence • Machine Learning • Software
Operate, scale, and optimize distributed storage infrastructure supporting large-scale AI/ML and HPC workloads. Build Python automation, manage Linux bare-metal systems, troubleshoot storage, hardware, networking, and operating system issues, and improve performance, reliability, monitoring, capacity planning, and lifecycle management. Collaborate with infrastructure, networking, platform, and data center teams on storage deployments and scaling strategies.
4 Hours AgoSaved
Remote or Hybrid
4 Locations
Artificial Intelligence • Machine Learning • Software
Build, validate, and operate large-scale bare-metal GPU infrastructure for AI/ML and HPC workloads. Responsibilities include managing Linux systems, image pipelines, test clusters, provisioning, firmware and driver validation, GPU diagnostics, performance analysis with NVIDIA DCGM, automation, virtualization, and hardware management interfaces. The role requires troubleshooting across hardware and software layers while collaborating with infrastructure, hardware, data center, platform, and ML teams.
4 Hours AgoSaved
Hybrid
San Francisco, CA, USA
Artificial Intelligence • Machine Learning • Software
Own revenue accounting and AR operations for a consumption-based GPU cloud business. Responsibilities include revenue close, ASC 606 technical accounting, CRM-to-cash data integrity, billing and collections, usage-to-cash reconciliations, contract cost accounting, audit support, historical revenue recasting, and building scalable CRM, billing, ERP, controls, and documentation processes.
4 Hours AgoSaved
Hybrid
San Francisco, CA, USA
Artificial Intelligence • Machine Learning • Software
Drives strategic finance, corporate development, and FP&A initiatives, including pricing, unit economics, GPU financing, M&A diligence, investor communications, and capital allocation. Builds and maintains financial models, leads reporting and forecasting cycles, supports annual and long-range planning, evaluates CapEx and infrastructure investments, and delivers executive-level insights on business performance, risks, and opportunities.
4 Hours AgoSaved
Hybrid
2 Locations
Artificial Intelligence • Machine Learning • Software
The Senior People Partner will support engineering and technical leaders on organizational design, leveling, promotions, compensation, employee relations, performance management, and exits. The role will strengthen manager capabilities, improve technical hiring practices, build scalable people systems, and manage global employment complexity. This position requires strong judgment, communication, and influence in a fast-scaling, distributed AI and technology company.
4 Hours AgoSaved
Hybrid
San Francisco, CA, USA
Artificial Intelligence • Machine Learning • Software
Designs and implements microservices-based platform services, RESTful APIs, backend systems, infrastructure automation, and cloud integrations. Responsibilities include architecture planning, monitoring and alerting, production troubleshooting, disaster recovery, performance optimization, code review, mentoring, and participation in technical planning. Requires strong experience with distributed systems, AWS, GCP or Azure, Kubernetes, microservices, programming in Go, Python or Java, infrastructure as code, networking, security, and scalability.
4 Hours AgoSaved
Remote or Hybrid
4 Locations
Artificial Intelligence • Machine Learning • Software
Operate and scale large GPU infrastructure platforms, including Linux systems, bare-metal environments, provisioning workflows, observability, and reliability automation. Responsibilities include platform deployment, incident response, break/fix operations, customer provisioning, on-call participation, infrastructure troubleshooting, and collaboration across engineering, networking, customer success, and software teams. The role also develops automation to reduce manual work and improve operational efficiency.
4 Hours AgoSaved
Hybrid
2 Locations
Artificial Intelligence • Machine Learning • Software
Own cross-functional operations projects that remove growth blockers, diagnose organizational friction, and build scalable AI-powered internal tools and automations. Partner with teams across the company, drive project adoption, maintain internal operations documentation, support business programs, and measure outcomes using data. The role requires strong analytical thinking, stakeholder management, end-to-end project ownership, and hands-on experience building workflows with APIs, SaaS platforms, automation tools, and AI coding tools.