Bright Vision Technologies Jobs

AI Performance Optimization Engineer

Bright Vision Technologies

AI Performance Optimization Engineer

Posted Yesterday

Be an Early Applicant

In-Office

San Ramon, CA, USA

100K-150K Annually

Senior level

In-Office

San Ramon, CA, USA

100K-150K Annually

Senior level

Optimize end-to-end training and inference for large neural networks by profiling pipelines, eliminating bottlenecks, implementing quantization/pruning, tuning distributed training (tensor/pipeline parallelism, FSDP/ZeRO), driving compiler/kernel optimizations (Triton/XLA/TVM), building benchmarks, and advising on hardware/software and cost-efficiency to translate research advances into production gains.

The summary above was generated by AI

Bright Vision Technologies is a forward-thinking software development company dedicated to building innovative solutions that help businesses automate and optimize their operations. We leverage cutting-edge technologies to create scalable, secure, and user-friendly applications.
As we continue to grow, we’re looking for a skilled AI Performance Optimization Engineer to join our dynamic team and contribute to our mission of transforming business processes through technology.
This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential.
Job Title: AI Performance Optimization Engineer
Location: 100% Remote (Continental United States)
Position Type: In-house Bright Vision Technologies SOW engagement (no third-party client or vendor)
Salary: $100K - $150K
Experience: 6+ years
Sponsorship: No new H1B sponsorship available. H1B transfers welcomed for qualified candidates.
Employment Type: Full-time, direct W2 with Bright Vision Technologies (no C2C, no 1099, no third-party)
Engagement: Long-term, multi-year, aligned to the Bright Vision SOW delivery roadmap
Compensation: Competitive base salary commensurate with experience, plus benefits.
Employment Terms & Visa Policy
This is a 100% remote, full-time, direct W2 position with Bright Vision Technologies.
This role is part of Bright Vision Technologies’ in-house Statement of Work (SOW) engagement. The client, end customer, and employer for this position is Bright Vision Technologies — there is no third-party client, vendor, or implementation partner involved.
We do not engage in C2C, 1099, or third-party arrangements for this role.
BUT STRICTLY NO C2C/1099/3RD PARTY COMPANIES. ALL OUR ROLES ARE W2 AND NO 3RD PARTY BROKERING PLEASE.
Candidates must be willing to work directly as a full-time W2 employee of Bright Vision Technologies and contribute to our in-house SOW deliverables.
No new H1B sponsorship is available for this role.
However, candidates who are currently on a valid H1B visa and require a transfer are welcome to apply. We will support H1B transfers for qualified candidates.
For every role, a technical coding assessment is mandatory. Please apply only if you are confident in your technical abilities and hands-on experience.
Job Summary
We are seeking an AI Performance Optimization Engineer to focus on extracting maximum throughput, minimizing latency, and reducing cost across training and inference workloads for large neural network systems. The role spans the full stack from low-level kernel optimization to distributed system tuning, requiring deep understanding of GPU architecture, model parallelism, memory management, and compiler-level optimization. The ideal candidate has demonstrated impact on production AI workloads, with strong instrumentation and measurement discipline that enables rigorous, data-driven optimization decisions. In this role you will work closely with cross-functional partners — product, design, engineering, operations, and business stakeholders — to translate ambiguous requirements into well-engineered solutions, and will be expected to raise the bar through code review, design review, and mentorship of more junior engineers. The successful candidate brings strong engineering discipline, a clear communication style, and a track record of shipping meaningful work that holds up well in production.
Key Responsibilities

Profile and optimize end-to-end AI training and inference pipelines for throughput, latency, and cost.
Identify and eliminate bottlenecks across data loading, model compute, communication, and memory.
Implement and tune quantization, sparsity, and pruning strategies to reduce model footprint and accelerate inference.
Optimize distributed training using tensor parallelism, pipeline parallelism, FSDP, and ZeRO-style sharding.
Tune attention implementations using FlashAttention, paged attention, and related techniques.
Implement KV cache optimization, continuous batching, and speculative decoding for LLM serving.
Drive compiler-level optimizations using Triton, XLA, TorchInductor, or TVM, working with the broader ML framework community to land improvements that translate into measurable end-to-end performance gains.
Optimize data pipelines, sharding strategies, and storage access patterns for high-throughput training.
Build and maintain rigorous benchmark suites and regression frameworks across workloads.
Collaborate with ML and platform engineering teams to embed best practices in standard pipelines.
Drive cost-efficiency improvements through model architecture, hardware selection, and scheduling strategies.
Evaluate new hardware and software offerings, and advise on adoption.
Document performance tuning playbooks and share findings broadly across engineering teams.
Stay current with AI systems research and translate advances into production improvements.

Required Qualifications

Bachelor’s or Master’s degree in Computer Science, Computer Engineering, or a related field.
Six or more years of experience in performance engineering, ML systems, or HPC.
Strong proficiency in Python and C++.
Hands-on experience optimizing deep learning workloads on modern GPUs.
Deep understanding of distributed training and inference techniques.
Experience with profiling tools across CPU, GPU, and distributed systems.
Familiarity with model compression techniques and their accuracy implications.
Strong grasp of memory hierarchies, communication primitives, and parallelism strategies.
Excellent measurement, debugging, and analytical reasoning skills.
Strong communication and collaboration skills.

Preferred Qualifications

Experience optimizing LLM inference at production scale.
Contributions to vLLM, TensorRT-LLM, DeepSpeed, or similar projects.
Familiarity with custom kernel authoring in Triton or CUTLASS.
Experience with FinOps for AI workloads.
Publications or talks on AI systems performance.

How to Apply
Would you like to know more about this opportunity?
For immediate consideration, please send your resume to [email protected]
Learn more about Bright Vision Technologies at www.bvteck.com.
We recognize that our people are our strength, and the diverse talents they bring to our global workforce are directly linked to our success. We are an equal opportunity employer and place a high value on diversity and inclusion at our company.
We do not discriminate on the basis of any protected attribute, including race, religion, color, national origin, gender, sexual orientation, gender identity, gender expression, age, marital or veteran status, pregnancy or disability, or any other basis protected under applicable law. We also make reasonable accommodations for applicants’ and employees’ religious practices and beliefs, as well as mental health or physical disability needs.
Bright Vision Technologies is an Equal Opportunity Employer, including Disability/Veterans.
Position offered by “No Fee Agency.”

Similar Jobs

Zyphra

Research Engineer - AI Performance & Kernel Optimization

15 Days Ago

In-Office

San Francisco, CA, USA

Mid level

Information Technology • Software

As a Research Engineer, you will optimize AI performance, focusing on kernel development for large-scale ML workloads, profiling bottlenecks, and collaborating with teams to enhance model training and inference.

Top Skills: AmdArmAws TrainiumCudaGoogle TpuHipHpcIntelPtxQualcommTriton

Grow Therapy

Staff Engineer

33 Minutes Ago

Remote or Hybrid

San Francisco, CA, USA

220K-240K Annually

Senior level

220K-240K Annually

Senior level

Healthtech • Social Impact • Software

Lead and execute Grow Therapy's multi-year Security Engineering roadmap. Build secure-by-default infrastructure (auth, authZ, logging, egress), drive data security and systematic data tagging, develop org-wide security scorecards, enable automated least-privilege and vulnerability management, and influence AI-native security and security culture across product, platform, and compliance teams.

PwC

Martech Developer- Manager

2 Hours Ago

Remote or Hybrid

212K-244K Annually

Mid level

212K-244K Annually

Mid level

Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI

Lead selection, implementation, and administration of marketing and sales technologies to drive growth and customer engagement. Manage and coach a team, execute digital marketing and creative campaigns, optimize marketing automation and Salesforce analytics, ensure data quality and validation, and partner with stakeholders to improve processes and deliverables from planning through completion.

Top Skills: Adobe Data CollectionAdobe Experience Manager (Aem)Adobe Martech PlatformsAnalytics InstrumentationCdpCRMDom ManipulationHTMLJavaScriptMarketing AutomationSalesforce Crm AnalyticsSalesforce Marketing CloudTypescriptWeb Sdk

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
Major Tech Employers: Google, Apple, Salesforce, Meta
Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine