CrowdStrike Logo

CrowdStrike

Sr. AI Infrastructure Engineer, LLM/AI Platforms (Remote)

Reposted 7 Days Ago
Remote or Hybrid
Hiring Remotely in USA
140K-215K Annually
Senior level
Remote or Hybrid
Hiring Remotely in USA
140K-215K Annually
Senior level
Design, build, and operate large-scale LLM infrastructure and data platforms for training, fine-tuning, and inference. Provision GPU clusters, optimize GPU utilization, implement model lifecycle management, deploy inference frameworks, create evaluation and observability systems, and collaborate with data scientists to productionize AI capabilities. Mentor engineers and enforce MLOps/DataOps best practices.
The summary above was generated by AI

As a global leader in cybersecurity, CrowdStrike protects the people, processes and technologies that drive modern organizations. Since 2011, our mission hasn’t changed — we’re here to stop breaches, and we’ve redefined modern security with the world’s most advanced AI-native platform. We work on large scale distributed systems, processing almost 3 trillion events per day and this traffic is growing daily. Our customers span all industries, and they count on CrowdStrike to keep their businesses running, their communities safe and their lives moving forward. We're proud to work for a mission-driven company leveraging AI to transform the way we work. CrowdStrikers drive their careers through flexibility and autonomy while also being expected to contribute to a culture of responsible AI adoption, experimentation, and innovation. We use an AI-first mindset as a force multiplier to proactively and continuously accelerate execution, build expertise, uncover insights, and solve complex problems. We’re always looking to add talented CrowdStrikers to the team who have limitless passion, a relentless focus on innovation and a fanatical commitment to our customers, our community and each other. Ready to join a mission that matters? The future of cybersecurity starts with you.


About the Role:

CrowdStrike is looking for a Senior AI Infrastructure Engineer with expertise in Large Language Models (LLMs) Infrastructure and data platforms to join our growing AI Infrastructure Team. You will be a key leader, helping to design, build, and deploy cutting-edge AI infrastructure that powers our next generation of AI-driven security products. This role requires hands-on experience in LLM infrastructure to support multiple large scale training pipelines and scalable AI-powered systems. You will champion engineering best practices, write high-quality code, and actively mentor and strengthen the team’s technical knowledge and capabilities.

CrowdStrike is a computer security company, but we do not require candidates for this role to have prior security industry experience. We will mentor and train in security topics as needed. We do expect a strong interest in CrowdStrike's mission and a willingness to engage with the needs of our product teams.

The scale of our systems and data are approaching Exabytes in size. Experience with extremely large-scale systems, including DevSecOps patterns, practices, and standards are important for this work.


What You'll Do:
  • Provision and configure large GPU clusters and compute resources for LLM training, finetuning, and inference workloads.

  • Develop and optimize LLM model-serving infrastructure, including deployment and optimization of various inference frameworks.

  • Lead model lifecycle management including versioning, checkpointing and reproducibility across training and inference deployments.

  • Design and champion robust evaluation frameworks to assess model performance, accuracy, and reliability, ensuring AI systems are consistently at production-ready standards.

  • Identify and address GPU utilization and GPU memory efficiency bottlenecks and apply techniques like quantization, batching, and caching.

  • Architect and maintain data platforms and pipelines specifically designed to support LLMs, Retrieval-Augmented Generation (RAG), and AI Agentic Systems at scale.

  • Deliver production-ready code with a focus on performance, maintainability, and testing rigor, ensuring the ability to ship fast without compromising quality.

  • Apply expertise in data modeling, normalization, and semantic cataloging for AI/ML workloads.

  • Define and enforce best practices for MLOps/DataOps surrounding LLMs, including monitoring, observability, and zero-touch recovery mechanisms for AI services.

  • Document architectural designs thoroughly and communicate technical decisions clearly to stakeholders

  • Collaborate across the organization with Data Scientists, Product Managers, and other engineering teams to transform research prototypes into robust, production-grade services.


Tech Stack (Experience in several areas is expected):
  • Hands-on experience with MLOps Tools (MLflow, Sagemaker, Vertex AI).

  • Strong understanding of CUDA, NVIDIA drivers, GPU, and TPU compute fundamentals.

  • Experience with inference serving frameworks such as vLLM and Triton Inference Server.

  • Proficiency with distributed training frameworks including Pytorch, Ray, Megatron, and JAX.

  • Expert-level proficiency in a high-level coding language (Python).

  • Deep knowledge of containerization and orchestration (Docker, Kubernetes, Slurm, Airflow).

  • Proficiency with Infrastructure as Code tooling like Terraform and Ansible.

  • Experience with cloud platforms (AWS, GCP, or OCI) and related data services.


What You'll Need:
  • Bachelor’s degree in Computer Science, Data Engineering, or a related STEM field; Master’s degree preferred

  • 6+ years of experience in Infrastructure/Data Engineering, with at least 2 years focused on building and maintaining platforms/pipelines that support LLM-based systems and applications

  • Demonstrable hands-on experience in LLM infrastructure engineering including cluster provisioning, optimizing training workloads, and maintaining inference pipelines

  • Exceptional ability to write clean, elegant, performant, and well-tested code, coupled with a strong focus on action and delivering results quickly.

  • Thorough understanding of engineering practices including effective peer code reviews and resilient architecture design

  • Demonstrates technical leadership and mentorship capabilities

  • Proven experience utilizing AI technologies to enhance decision-making, streamline workflows and processes, improve efficiency and drive business outcomes.

Bonus Points:

  • Prior experience in the cybersecurity, intelligence, or high-compliance industries.

  • Direct experience building, deploying, and managing LLMs in a production environment.

  • Experience with common agentic workflow frameworks (e.g., LangChain, LlamaIndex).

  • Experience with distributed data processing frameworks (e.g., Spark, Dask, Flink).

#LI-DM1

#LI-Remote

Benefits of Working at CrowdStrike:

  • Market leader in compensation and equity awards

  • Comprehensive physical and mental wellness programs 

  • Competitive vacation and holidays for recharge  

  • Paid parental and adoption leaves

  • Professional development opportunities for all employees regardless of level or role

  • Employee Networks, geographic neighborhood groups, and volunteer opportunities to build connections

  • Vibrant office culture with world class amenities

  • Great Place to Work Certified™ across the globe

CrowdStrike is proud to be an equal opportunity employer. We are committed to fostering a culture of belonging where everyone is valued for who they are and empowered to succeed. We support veterans and individuals with disabilities through our affirmative action program.

CrowdStrike is committed to providing equal employment opportunity for all employees and applicants for employment. The Company does not discriminate in employment opportunities or practices on the basis of race, color, creed, ethnicity, religion, sex (including pregnancy or pregnancy-related medical conditions), sexual orientation, gender identity, marital or family status, veteran status, age, national origin, ancestry, physical disability (including HIV and AIDS), mental disability, medical condition, genetic information, membership or activity in a local human rights commission, status with regard to public assistance, or any other characteristic protected by law. We base all employment decisions--including recruitment, selection, training, compensation, benefits, discipline, promotions, transfers, lay-offs, return from lay-off, terminations and social/recreational programs--on valid job requirements.

If you need assistance accessing or reviewing the information on this website or need help submitting an application for employment or requesting an accommodation, please contact us at [email protected] for further assistance.

Find out more about your rights as an applicant.

CrowdStrike participates in the E-Verify program.

Notice of E-Verify Participation

Right to Work

CrowdStrike, Inc. is committed to fair and equitable compensation practices. Placement within the pay range is dependent on a variety of factors including, but not limited to, relevant work experience, skills, certifications, job level, supervisory status, and location. The base salary range for this position for all U.S. candidates is $140,000 - $215,000 per year, with eligibility for bonuses, equity grants and a comprehensive benefits package that includes health insurance, 401k and paid time off.

For detailed information about the U.S. benefits package, please click here

Expected Close Date of Job Posting is:09-06-2026

CrowdStrike Sunnyvale, California, USA Office

150 Mathilda Place, Sunnyvale, CA, United States, 94086

Similar Jobs at CrowdStrike

Yesterday
Remote or Hybrid
USA
85K-120K Annually
Senior level
85K-120K Annually
Senior level
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Build and operate cloud and on-premises network infrastructure supporting CrowdStrike’s global security platform. Responsibilities include developing Terraform modules, implementing AWS multi-region and hybrid networking architectures, supporting cloud migration and data center exit programs, contributing to architecture reviews, improving CI/CD and Kubernetes enablement, documenting standards, and monitoring infrastructure. The role requires cross-functional communication, compliance-aware design, infrastructure automation, and use of AI technologies to improve troubleshooting and operational efficiency.
Top Skills: ArgocdAWSAws Network FirewallCiscoDatadogGitlab Ci/CdGitopsHelmIamKubernetesOctopus DeployTerraformTransit GatewayVpc
Yesterday
Remote or Hybrid
CA, USA
125K-180K Annually
Senior level
125K-180K Annually
Senior level
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Lead hands-on engineering delivery of agentic AI solutions for GTM systems. Design and build LLM-powered workflows, autonomous agents, RAG and semantic search pipelines, Salesforce and Slack integrations, CI/CD and observability, enforce AI governance/security, mentor engineers, and drive prototypes to production across enterprise GTM platforms.
Top Skills: AgentcoreAgentforceApexAutogenAws BedrockCi/CdCopadoCrewaiGithub ActionsJavaScriptJenkinsLangchainLanggraphLightning Web ComponentsLlamaindexLwcMcpPythonRag (Retrieval Augmented Generation)RestSalesforce Platform EventsSemantic KernelSemantic SearchSlackSlack Workflow BuilderSoapTypescriptVector DatabasesVertex AiWorkflow Orchestration
Yesterday
Remote or Hybrid
CA, USA
85K-128K Annually
Mid level
85K-128K Annually
Mid level
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Pre-sales specialist for NG Identity Security responsible for driving sales in California, delivering technical demos and proofs of value, advising customers, collaborating with sales, product and engineering teams, shaping go-to-market strategy, and expanding product adoption.
Top Skills: Adaptive ShieldAICrowdstrike FalconMicrosoft 365SalesforceSspm

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account