Bright Vision Technologies Logo

Bright Vision Technologies

Reliability Engineer

Posted 15 Days Ago
Be an Early Applicant
In-Office
2 Locations
75K-95K Annually
Senior level
In-Office
2 Locations
75K-95K Annually
Senior level
Operate and improve the reliability, availability, and performance of large-scale distributed systems. Build automation and tooling, manage Linux and Kubernetes environments, design CI/CD pipelines, implement observability, lead incident response, conduct post-incident reviews, and apply SLOs, error budgets, capacity planning, chaos engineering, and cloud platform practices.
The summary above was generated by AI
Reliability Engineer – Remote
Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.
This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential.
Job Title: Reliability Engineer
Location: 100% Remote (U.S.)
Position Type: Full-time, Direct W2
Salary Range: $75,000–$95,000 Annually
Experience Required: 6+ years
Sponsorship: U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position.
Job Summary
We are seeking an experienced Reliability Engineer to ensure the availability, performance, and operational excellence of large-scale distributed systems in production. As an SRE you will live at the boundary between development and operations, applying strong software engineering principles to infrastructure and operations problems, and continually pushing the platform toward higher reliability with lower operational toil. The ideal candidate will combine deep systems knowledge with strong programming skills, a measurement-driven mindset, and the discipline to design, automate, and operate complex services so that reliability becomes a first-class engineering deliverable rather than a reactive concern.
Required Qualifications
  • Bachelor’s degree in Computer Science, Engineering, or a related technical discipline.
  • Five or more years of SRE, DevOps, or production engineering experience supporting large-scale distributed systems.
  • Strong programming skills in at least one of Python, Go, or Java, with the ability to build robust automation and tooling.
  • Deep, hands-on experience operating Linux at scale, including networking, performance tuning, and systems-level troubleshooting.
  • Production experience operating Kubernetes and container-based workloads.
  • Strong working knowledge of observability tooling such as Prometheus, Grafana, OpenTelemetry, ELK/EFK, or commercial equivalents.
  • Hands-on experience designing and operating CI/CD pipelines for both infrastructure and applications.
  • Solid understanding of distributed system design, including consistency models, partitioning, and failure semantics.
  • Demonstrated experience leading incident response and conducting effective post-incident reviews.
  • Excellent communication and documentation skills.
Preferred Qualifications
  • Experience defining and operationalizing SLOs and error budgets in real production environments.
  • Exposure to chaos engineering practices and tools such as Chaos Monkey, Gremlin, or Litmus.
  • Hands-on experience with at least one major cloud platform (AWS, Azure, or GCP).
  • Background in capacity planning, performance engineering, or large-scale load testing.
  • Familiarity with service mesh technologies such as Istio, Linkerd, or Consul.
How to Apply
Would you like to know more about this opportunity? For immediate consideration, please send your resume to [email protected]. Learn more about Bright Vision Technologies at www.bvteck.com.
Bright Vision Technologies is an Equal Opportunity Employer.

Similar Jobs

7 Days Ago
Easy Apply
Remote or Hybrid
USA
Easy Apply
225K-265K Annually
Senior level
225K-265K Annually
Senior level
Fintech • Information Technology • Software • Financial Services
Own the durability, recoverability, performance, and security of a production PostgreSQL/RDS fleet supporting a live trading platform. Lead replication, failover, backup and restore, disaster-recovery drills, data lifecycle management, access control, encryption, and database observability. Investigate engine-level performance issues including WAL contention, replica lag, bloat, and locking. Build infrastructure and AI-assisted operational tooling while documenting runbooks and reliability decisions.
Top Skills: AlloydbAmazon AuroraAmazon RdsAWSBashBigQueryElkGrafanaKafkaKubernetesLinuxPostgresPrometheusPythonSQLTerraform
7 Days Ago
In-Office
138K-230K Annually
Senior level
138K-230K Annually
Senior level
3D Printing • Aerospace • Hardware • Software • Manufacturing
Own build reliability for flight avionics hardware across its lifecycle. Lead readiness criteria, manufacturing process maturation, quality reviews, risk retirement, root-cause investigations, containment, corrective actions, cost reduction, and cross-functional alignment across engineering, manufacturing, quality, supply chain, and mission assurance. Analyze build data, improve process controls, support mission-critical events, and mentor teams while ensuring reliable production and flight readiness for human-rated space station systems.
Top Skills: As9100AvionicsControl PlansFailure Mode And Effects Analysis (Fmea)Iso 9001Root Cause Analysis (Rca)Six SigmaStatistical Process Control (Spc)
7 Days Ago
In-Office
162K-265K Annually
Senior level
162K-265K Annually
Senior level
3D Printing • Aerospace • Hardware • Software • Manufacturing
Leads reliability engineering for Haven-1 avionics, RF, and power systems. Responsibilities include FMEA and fault-tree analysis, design and test reviews, verification and validation, reliability risk reduction, engineering standards, configuration management, change control, and hardware integration. The role guides design decisions, qualification and acceptance testing, manufacturing readiness, and human-rated space station reliability throughout the development lifecycle.
Top Skills: AvionicsConfiguration ManagementEmi/EmcFault Tree AnalysisFmeaMatlabExcelPower SystemsPythonRfSystems IntegrationTableauVerification And Validation

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account