Provides day-to-day production support for complex cloud application systems. Responsibilities include outage detection and resolution, system administration, infrastructure maintenance, user support, monitoring, automation, database support, asset inventory, security compliance, documentation, and continuous improvement. The role collaborates with development, operations, security, customers, and management, supports AWS and Azure environments, and participates in on-call production support.
SAIC is searching for a Systems Engineer who performs high-level, day-to-day operational support of complex application cloud systems to join our VA team. Develops solutions to routine technical problems of limited scope. Follows standard practices and procedures in analyzing situations or data from which answers can be readily obtained. This is a 100% remote role.
Job Description:
- Detect, isolate, document, rapidly report, and resolve system outages or problems encountered during operations of the scientific workstations, which includes the collections of diagnostic data, restoring the system operation, development of workarounds, and other activities necessary for recovery of a system.
- Accurately document problems in logging and discrepancy reporting tools.
- Work directly with the customer in most aspects of the day-to-day activities.
- Respond to user calls regarding hardware and software problems, correcting or ensuring that problems are escalated when required. Communicate with users and senior management the status of key problem statuses.
- Perform maintenance/installation of computing infrastructure.
- Implement continuous improvement methodology through the use of IT systems or procedure.
- Maintain inventory of system assets.
- Ensure compliance with VA standards and security policies.
- Provide documentation, training, and additional duties as assigned.
Education:
- Bachelor’s or Master’s degree in Computer Science, Information Technology, Systems Engineering, or a related field.
- Advanced certifications in relevant areas (e.g., Red Hat Certified Engineer, Microsoft Certified Systems Engineer, AWS Certified Solutions Architect) are preferred.
Experience:
- 8 years of experience in systems engineering or a related field, particularly in production environments.
- Experience with AWS and/or Azure Cloud experience with S3, Step Functions, Batch Jobs, CloudWatch.
- Minimum 3 years SQL query and monitoring experience.
- Proven track record of managing and maintaining large-scale production systems.
Requirements:
- Strong proficiency in Linux/Unix and Windows operating systems.
- Experience with system administration tasks including user management, permissions, and system monitoring.
- Proficiency in scripting languages such as Python, Bash, or PowerShell for automation and configuration management.
- Experience with automation tools like Ansible, Puppet, Chef, or SaltStack.
- Knowledge of cloud-native technologies and infrastructure-as-code (IaC) tools such CloudFormation.
- Experience with monitoring tools like DynaTrace, ScienceLogic, and CloudWatch.
- Ability to troubleshoot and optimize system performance and reliability.
- Understanding of network protocols, firewall configurations, and VPN setup.
- Experience with network monitoring and diagnostic tools.
- Experience with security best practices for production systems.
- Experience implementing security measures, conducting audits, and ensuring compliance with industry standards.
- Working knowledge of database systems such as AWS RDS SQL Server and Oracle RDMBS/RDS.
- Knowledge of database performance tuning and backup/recovery processes.
- Knowledge of continuous integration/continuous deployment (CI/CD) pipelines.
- Familiarity with version control systems like Git and CI/CD tools like GitHub Actions, AWS CodeBuild and CodeDeploy.
- Strong analytical and problem-solving skills with the ability to troubleshoot complex system issues.
- Excellent communication and interpersonal skills, with the ability to work effectively in cross-functional teams.
- Ability to prioritize tasks, manage time effectively, and meet deadlines in a high-pressure environment.
- Strong attention to detail to ensure system stability and data integrity.
- Ability to quickly adapt to new technologies and processes.
- Willingness to be on-call for production system support as required.
- Strong documentation skills for maintaining system configurations, processes, and procedures.
- Ability to manage and contribute to multiple projects, ensuring timely completion and quality results.
- Ability to work collaboratively with development, operations, and security teams to ensure seamless production system operations.
We are approximately 23,000 strong; driven by mission, united by purpose, and inspired by opportunities. SAIC is an Equal Opportunity Employer. Headquartered in Reston, Virginia, SAIC has annual revenues of approximately $7.3 billion. For more information, visit saic.com. For ongoing news, please visit our newsroom.
Spectrum San Diego Richmond, California, USA Office
Richmond, United States
Similar Jobs
Security
Provide day-to-day production support for complex cloud systems and scientific workstations. Diagnose and resolve outages, maintain infrastructure, manage system assets, support users, monitor performance, implement automation and continuous improvement, and ensure compliance with VA security standards. Responsibilities include system administration, cloud and database operations, documentation, training, troubleshooting, and collaboration with development, operations, and security teams. Participation in on-call production support is required.
Top Skills:
Amazon CloudwatchAmazon S3AnsibleAWSAws BatchAws CloudformationAws CodebuildAws CodedeployAws RdsAws Step FunctionsAzureBashChefCi/CdDynatraceGitGithub ActionsLinux/UnixOracle RdbmsPowershellPuppetPythonSaltstackSciencelogicSQLSQL ServerWindows
Financial Services
Leads software engineering and AI/ML solution delivery using Java and Python. Designs, builds, operates, and troubleshoots secure, scalable production systems; advances AI-assisted engineering practices; integrates LLM, RAG, and NLP capabilities; develops CI/CD and cloud-native automation; improves observability, reliability, and incident remediation; partners on technical roadmaps; and mentors engineers and AI practitioners.
Top Skills:
Ai/MlAWSCi/CdCloud-Native ArchitecturesContainerizationEvent-Driven ArchitectureGenerative AiInfrastructure As CodeJavaLangchainLanggraphLlmsMicroservicesNlpObservabilityPythonRetrieval-Augmented Generation (Rag)
Aerospace • Hardware • Information Technology • Security • Software • Cybersecurity • Defense
Develop and maintain Master Data Management, data governance, stewardship, quality, and reference data processes. Partner with enterprise stakeholders and governance groups to define standards, validate master data, establish quality rules, support data owners, and prioritize initiatives. Lead requirements for system configuration changes while using data profiling, modeling, and analytics to improve data validity and consistency across systems.
Top Skills:
AgileInformaticaProfiseePythonRestSemarchySoapSQL
What you need to know about the San Francisco Tech Scene
San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.
Key Facts About San Francisco Tech
- Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
- Major Tech Employers: Google, Apple, Salesforce, Meta
- Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
- Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
- Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
- Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine


