People Finders Logo

People Finders

Director, Infrastructure & IT Operations

Posted 3 Days Ago
Remote
Hiring Remotely in USA
170K-180K Annually
Expert/Leader
Remote
Hiring Remotely in USA
170K-180K Annually
Expert/Leader
Leads cloud infrastructure, SRE, production operations, corporate IT, helpdesk, and bot mitigation. Owns AWS architecture, reliability, security, observability, incident response, disaster recovery, identity management, vendor relationships, budgets, and cloud costs. Builds operational standards, SLOs, automation, infrastructure roadmaps, and high-performing teams while improving customer-facing platform availability, traffic protection, employee technology support, and cybersecurity readiness.
The summary above was generated by AI

PeopleFinders.com, the premier online service for consumers to locate, contact and verify people and businesses. Over the past couple of decades the Company has quietly become one of the largest owners of public records data in the country, distributing its products over a vast network of websites.

PeopleFinders | Remote |

Salary - $170k - $180k

About the Role

We are seeking a strategic and hands-on Director, Infrastructure & IT Operations to lead the technology capabilities that keep our digital products, cloud platforms, and employees secure, reliable, and productive.

This leader will oversee cloud infrastructure, Site Reliability Engineering, production operations, bot mitigation, corporate IT, and helpdesk support, while providing architectural leadership across cloud platforms, networking, identity, observability, and operational tooling.

The ideal candidate combines strong technical judgment with disciplined operational leadership: building high-performing teams, setting clear service expectations, improving reliability, managing vendors and costs, and reducing operational risk across the company's portfolio of digital properties.

This is not a traditional corporate IT position — it spans both internal employee technology and the production infrastructure supporting high-traffic, customer-facing applications.

Key Responsibilities

Infrastructure and Technology Leadership

Lead the teams responsible for cloud infrastructure, SRE, production operations, corporate IT, helpdesk, and bot mitigation. Establish ownership, operational standards, SLOs, escalation procedures, and performance metrics for each function. Develop managers, technical leads, and IT staff through coaching and career development. Build quarterly and annual infrastructure roadmaps aligned with business priorities, growth, security, reliability, and cost. Partner with Engineering, Product, Data, Marketing, Finance, HR, Legal, and executives, providing clear recommendations on risks, architecture, and tradeoffs. Build a culture of accountability, documentation, automation, and continuous improvement.

Cloud Infrastructure and Architecture

Own the reliability, scalability, security, performance, and cost effectiveness of the company's cloud infrastructure, primarily AWS. Provide architectural leadership across cloud services, applications, APIs, networking, identity, databases, storage, observability, and integrations. Establish architecture standards, reference patterns, governance, and technical review processes. Partner with engineering to improve deployment safety, performance, and readiness. Drive automation, configuration management, and infrastructure as code. Lead capacity planning for business growth and demand spikes. Eliminate single points of failure, undocumented systems, manual processes, and reliance on individual employees or vendors. Review new systems for architectural fit, supportability, security, and total cost of ownership. Ensure production and nonproduction environments are properly separated, secured, and monitored.

Site Reliability and Production Operations

Establish and maintain SLIs, SLOs, availability targets, and error-budget practices for critical systems. Improve observability through centralized logging, metrics, distributed tracing, APM, and actionable alerting. Lead major incident management, including coordination, executive communication, root-cause analysis, and corrective-action tracking. Develop and test backup, disaster recovery, and business continuity procedures. Improve mean time to detect, respond, and recover from incidents. Establish on-call practices balancing coverage with team sustainability. Partner with development teams to improve resilience and production support. Use incident data to prioritize long-term reliability work. Ensure systems have current runbooks, architecture diagrams, and documentation.

Bot Mitigation and Traffic Protection

Own the strategy for bot mitigation, scraping protection, automated abuse prevention, and traffic-quality management. Protect company properties from scraping, credential attacks, fraud, automated abuse, and infrastructure exhaustion. Lead configuration of technologies such as Cloudflare, DataDome, WAFs, rate limiting, CAPTCHA alternatives, device fingerprinting, and application-level defenses. Partner with Engineering, Product, Marketing, SEO, and Revenue to distinguish malicious automation from legitimate customers, search engines, partners, and approved AI crawlers. Establish monitoring and response procedures for scraping and abnormal traffic. Measure effectiveness through traffic reduction, false-positive rates, platform stability, and infrastructure savings. Evaluate vendors for measurable value, and stay current on scraping techniques, browser automation, residential proxies, and AI crawlers.

Corporate IT and Helpdesk

Oversee internal IT services and helpdesk support for employees. Establish service-level expectations for response time, resolution time, satisfaction, and backlog reduction. Own the employee technology lifecycle: onboarding, offboarding, equipment provisioning, access management, and asset recovery. Manage laptops, endpoints, collaboration platforms, productivity software, and telecom services. Standardize endpoint configuration, encryption, patching, and endpoint detection. Improve the support experience through clear intake, self-service resources, and automation. Maintain accurate inventories of devices, licenses, and assets. Partner with HR so new employees get equipment and access on time and departing employees have access promptly removed. Identify and address the root causes of recurring support issues.

Identity, Access, and Operational Security

Establish consistent identity and access-management practices across corporate and production systems. Implement role-based access, least-privilege principles, MFA, periodic access reviews, and timely access removal. Maintain secure processes for privileged, administrative, and service accounts. Partner with security and legal to identify and reduce technology risks. Support vulnerability remediation, security assessments, audits, and vendor reviews. Ensure cloud environments, endpoints, applications, and third-party services meet company security standards, and maintain readiness for cybersecurity incidents and applicable privacy requirements.

Vendor, Budget, and Cost Management

Manage infrastructure, security, IT, telecom, monitoring, and support vendors, including evaluations, contract reviews, renewals, licensing, and SLAs. Develop and manage infrastructure and IT budgets. Improve cloud cost visibility through tagging, allocation, and forecasting. Identify opportunities to eliminate unused services, consolidate tools, and renegotiate contracts. Communicate budget performance and investment recommendations to leadership. Ensure critical vendors have appropriate security controls and documented ownership, and reduce dependency on vendors through internal knowledge.

Leadership Expectations

Operate as both a strategic technology leader and an effective technical decision-maker. Remain calm, organized, and decisive during outages, security events, and high-pressure situations. Create accountability without unnecessary bureaucracy. Communicate technical risks and recommendations clearly to technical and nontechnical audiences. Build productive partnerships across Engineering, Product, Data, Marketing, Security, and business operations. Make decisions based on reliability, risk, customer impact, cost, and measurable outcomes. Develop leaders who can independently manage their functions while maintaining consistent standards, and encourage teams to solve root causes rather than rely on repeated manual intervention.

Required Qualifications

10+ years in cloud infrastructure, SRE, platform engineering, DevOps, IT operations, or related disciplines. 5+ years leading engineering or technology teams, including managers or technical leads. Demonstrated experience operating high-traffic, customer-facing platforms in AWS or a comparable cloud. Strong understanding of cloud architecture, networking, identity, security, observability, databases, APIs, storage, and distributed systems. Experience managing production availability, incident response, root-cause analysis, disaster recovery, and business continuity. Experience with infrastructure automation, configuration management, CI/CD, and infrastructure as code. Experience managing corporate IT, employee support, endpoint management, licensing, and access provisioning. Experience establishing operational metrics, SLOs, and support standards. Demonstrated ability to manage vendors, contracts, budgets, and cloud costs. Strong written and verbal communication skills, with the ability to explain technical issues and tradeoffs to executives.

Preferred Qualifications

Experience with Cloudflare, DataDome, or comparable bot-management and web application protection platforms, and leading bot mitigation or traffic-quality programs. Experience with AWS services, containerized environments, infrastructure as code, centralized logging, and APM (e.g., AppDynamics). Experience supporting large-scale consumer subscription, public-record, data-as-a-service, advertising, or ecommerce platforms. Experience modernizing legacy infrastructure, transitioning systems from external vendors, and consolidating cloud accounts or IT services. Familiarity with privacy requirements and secure handling of sensitive consumer data. Experience building or maturing SRE, DevOps, or platform engineering functions.

Measures of Success

Success will be measured through: improved availability and stability of customer-facing platforms; reduced impact from bots, scraping, and abuse; faster incident detection and recovery; fewer recurring incidents; improved observability and reduced alert noise; tested backup and disaster recovery capabilities; lower cloud waste and better cost visibility; consistent helpdesk performance; reliable onboarding, offboarding, and access management; stronger endpoint security and asset management; reduced reliance on undocumented systems and vendors; stronger architecture governance; and improved team accountability and development

Benefits
401k
401k match
Medical/Dental/Vision/ Life Insurance


Similar Jobs

47 Minutes Ago
Remote or Hybrid
129K-233K Annually
Mid level
129K-233K Annually
Mid level
eCommerce • Fintech • Hardware • Payments • Software • Financial Services
Own and grow Square’s Detroit territory through field-based prospecting, business visits, live product demonstrations, consultative selling, and full-cycle deal closing. Build pipeline through cold outreach, networking, events, referrals, and partnerships; develop relationships with local businesses; support onboarding; maintain Salesforce activity and forecasts; and consistently exceed sales quotas across Square’s software, hardware, and financial services products.
Top Skills: Payment Processing TechnologySalesforceSquare
47 Minutes Ago
Remote or Hybrid
7 Locations
164K-297K Annually
Senior level
164K-297K Annually
Senior level
eCommerce • Fintech • Hardware • Payments • Software • Financial Services
Own the full outbound sales cycle for mid-market merchants, from prospecting and pipeline development through discovery, product demonstrations, negotiation, and close. Build net-new business through strategic outreach, sell multi-product solutions, manage complex multi-stakeholder deals, forecast accurately in Salesforce, and consistently exceed revenue targets. Collaborate with business development, product, marketing, implementation, and operations teams while serving as a consultative advisor to merchants.
Top Skills: Salesforce
47 Minutes Ago
Remote or Hybrid
7 Locations
95K-168K Annually
Entry level
95K-168K Annually
Entry level
eCommerce • Fintech • Hardware • Payments • Software • Financial Services
Develop data science solutions for payments across Square and Cash App. Build AI-powered tools, ETL pipelines, dashboards, machine learning models, experiments, and cloud data systems. Partner with Product, Engineering, Finance, and Operations to improve payment success, cost efficiency, risk mitigation, and decision-making. Translate business needs into practical solutions, research questions independently, and operationalize data processes and infrastructure.
Top Skills: AICloud InfrastructureETLMachine LearningPythonSQL

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account