Xbox Logo

Xbox

Senior Data Ops Engineer, Data Activation & Products - Activision

Posted 9 Days Ago
Be an Early Applicant
In-Office
Santa Monica, CA
103K-190K Annually
Senior level
In-Office
Santa Monica, CA
103K-190K Annually
Senior level
Operate and improve the reliability, observability, and deployment safety of cloud, Kubernetes, application, and data-platform environments. Responsibilities include incident response, troubleshooting Databricks and Spark workloads, maintaining dashboards and alerts, supporting streaming and orchestration systems, automating operational tasks, improving release processes, and partnering with engineering teams on platform modernization and post-incident corrective actions.
The summary above was generated by AI
Job Title:Senior Data Ops Engineer, Data Activation & Products - Activision

Requisition ID:R028025

Job Description:

Your Mission

We are looking for a Site Reliability Engineer to help improve the reliability, observability, and operational maturity of our data platforms, Kubernetes-based deployment systems, internal applications, and cloud environments.

This role sits at the intersection of SRE, DevOps, and data. The ideal candidate is comfortable operating production systems, troubleshooting across infrastructure and applications, and helping teams deploy and support services more safely.

The role does not require someone to be a data engineer, but they should be excited about the systems that support modern data engineering, including Databricks, Spark, Airflow/Astronomer, streaming pipelines, event systems, internal tools, and backend services.

We are especially interested in a creative, curious engineer who enjoys learning new technology and using AI-accelerated development practices to solve problems faster and more thoughtfully. You should be excited to experiment with agentic development tools, automation frameworks, and emerging platform capabilities, while applying sound engineering judgment.

This role has the option to be based in our Los Angeles (Pen Factory) office and follows an onsite work schedule of Monday through Thursday or be remote. Work arrangements may change at the company’s discretion to meet business needs.

Priorities can often change in a fast-paced environment like ours, so this role includes, but is not limited to, the following responsibilities:    

  • Monitor service health, respond to alerts, and participate in incident response for cloud, Kubernetes, application, and data platform environments.
  • Investigate reliability issues across Kubernetes, networking, DNS, application runtime behavior, Databricks jobs, Spark workloads, orchestration systems, event systems, and dependent services.
  • Support the reliability of internal applications, APIs, workers, streaming consumers, event-driven services, and deployment workflows used by data engineering and business teams.
  • Build and maintain dashboards, alerting, runbooks, and operational documentation that improve detection and recovery speed.
  • Improve observability for Databricks environments, including job health, Spark streaming workloads, structured streaming metrics, cluster behavior, failures, latency, throughput, and cost signals.
  • Help route Spark streaming metrics, operational logs, event-system signals, and platform health signals into monitoring tools such as Grafana.
  • Contribute to alerting patterns for Databricks workflows, Airflow/Astronomer DAGs, dbt jobs, data freshness, pipeline failures, event lag, dead-letter queues, and production data dependencies.
  • Contribute scripts and automation that reduce repetitive operational work and improve environment hygiene.
  • Support release and deployment reliability by validating changes, improving rollback readiness, and strengthening change safety.
  • Partner with data engineers, analytics engineers, and software engineers to improve reliability across pipelines, services, internal tools, event systems, and data products.
  • Participate in post-incident follow-up and help close corrective actions that prevent recurrence.
  • Support platform modernization and migration efforts, including orchestration platform changes, deployment system improvements, and shared reliability standards.

Qualifications

  • 5+ years of experience in SRE, DevOps, cloud infrastructure, platform engineering, software engineering, data platform operations, or related production-support roles.
  • Hands-on experience supporting Kubernetes-based workloads, deployment systems, cloud infrastructure, or production application environments.
  • Familiarity with Linux, HTTP, DNS, containers, Kubernetes, Git-based workflows, and scripting in Bash, Python, or similar languages.
  • Experience with monitoring, logs, metrics, dashboards, alerting, and incident management practices.
  • Comfort working with event systems such as Kafka, Google Pub/Sub, Kinesis, or similar technologies, including topics, subscriptions, consumers, retries, lag, and dead-letter queues.
  • Strong troubleshooting mindset, clear communication, and comfort operating in a production-support environment.
  • Interest in data platforms, data engineering systems, orchestration, streaming workloads, event-driven architecture, internal developer tools, and production data services.

Nice to Have

  • Experience with Kubernetes deployment and release tooling such as Helm, ArgoCD, or similar GitOps workflows.
  • Experience with CI/CD automation using GitHub Actions, GitLab CI, Jenkins, or similar pipelines.
  • Familiarity with infrastructure as code using Terraform or similar provisioning tools.
  • Familiarity with Databricks, Spark, Spark Structured Streaming, Airflow, Astronomer, dbt, Kafka, Pub/Sub, object storage, or lakehouse architectures.
  • Experience with observability tools such as Grafana, Prometheus, Cloud Monitoring, Datadog, Splunk, or similar platforms.
  • Experience supporting internal applications, APIs, event-driven services, streaming consumers, or backend workers.
  • Exposure to data reliability concepts such as freshness, latency, completeness, pipeline health, data quality checks, and dependency-aware alerting.
  • Exposure to SLOs, SLAs, error budgets, postmortems, or formal reliability practices.

What Success Looks Like

  • Kubernetes and deployment workflows are easier to operate, monitor, and troubleshoot.
  • Databricks, Spark, streaming, orchestration, event-driven, and dbt environments have clearer dashboards, alerts, and runbooks.
  • Spark streaming metrics, event-system health signals, and platform logs are easier to access, visualize, and operationalize through tools such as Grafana.
  • Incidents are detected faster, resolved more efficiently, and followed up with meaningful corrective actions.
  • Data engineers, analytics engineers, and software engineers can deploy changes more safely and with greater confidence.
  • Repetitive operational work is reduced through automation and better platform standards.
  • Platform migrations and modernization efforts are supported with strong reliability, observability, and operational practices.

Our World 

At Activision, we strive to create the most iconic brands in gaming and entertainment. We’re driven by our mission to deliver unrivaled gaming experiences for the world to enjoy, together. We are home to some of the most beloved entertainment franchises including Call of Duty®, Crash Bandicoot™, Tony Hawk’s™ Pro Skater™, and Guitar Hero®. As a leading worldwide developer, publisher and distributor of interactive entertainment and products, our “press start” is simple: delight hundreds of millions of players around the world with innovative, fun, thrilling, and engaging entertainment experiences.

We’re not just looking back at our decades-long legacy; we’re forging ahead to keep advancing gameplay with some of the most popular titles and sophisticated technology in the world. We have bold ambitions to create the most inclusive company as we know our success comes from the passionate, creative, and diverse teams within our organization. 

We’re in the business of delivering fun and unforgettable entertainment for our player community to enjoy. And our future opportunities have never been greater — this could be your opportunity to level up. 

Ready to Activate Your Future? 

Rewards

We provide a suite of benefits that promote physical, emotional and financial well-being for ‘Every World’ - we’ve got our employees covered!  Subject to eligibility requirements, the Company offers comprehensive benefits including:

  • Medical, dental, vision, health savings account or health reimbursement account, healthcare spending accounts, dependent care spending accounts, life and AD&D insurance, disability insurance;
  • 401(k) with Company match, tuition reimbursement, charitable donation matching;
  • Paid holidays and vacation, paid sick time, floating holidays, compassion and bereavement leaves, parental leave;
  • Mental health & wellbeing programs, fitness programs, free and discounted games, and a variety of other voluntary benefit programs like supplemental life & disability, legal service, ID protection, rental insurance, and others;
  • If the Company requires that you move geographic locations for the job, then you may also be eligible for relocation assistance.

Eligibility to participate in these benefits may vary for part time and temporary full-time employees and interns with the Company.  You can learn more by visiting https://www.benefitsforeveryworld.com/.

In the U.S., the standard base pay range for this role is $102,800.00 - $190,204.00 Annual. These values reflect the expected base pay range of new hires across all U.S. locations. Ultimately, your specific range and offer will be based on several factors, including relevant experience, performance, and work location. Your Talent Professional can share this role’s range details for your local geography during the hiring process. In addition to a competitive base pay, employees in this role may be eligible for incentive compensation. Incentive compensation is not guaranteed. While we strive to provide competitive offers to successful candidates, new hire compensation is negotiable.

Similar Jobs

27 Minutes Ago
Hybrid
15-18 Hourly
Entry level
15-18 Hourly
Entry level
eCommerce • Fashion • Retail • Sales • Wearables • Design
The Barista II delivers exceptional customer experiences, engages guests, drives sales, maintains operational standards, and collaborates with team members.
An Hour Ago
Remote or Hybrid
2 Locations
286K-392K Annually
Senior level
286K-392K Annually
Senior level
Fintech • Machine Learning • Payments • Software • Financial Services
Designs, develops, deploys, and supports large-scale AI systems, including foundation models, LLM inference, agentic workflows, similarity search, guardrails, and model evaluation. Defines enterprise AI architecture, optimizes model performance, cost, latency, and throughput, establishes AI safety and governance standards, leads multi-year platform initiatives, and mentors senior technical leaders across engineering and research.
Top Skills: Agentic AiAi GovernanceAi ObservabilityAWSAws UltraclustersAzureC#C++CudaFoundation ModelsGoGCPHugging FaceJavaLarge Language ModelsModel EvaluationMulti-Agent WorkflowsPythonPyTorchScalaSimilarity SearchVectordbs
An Hour Ago
In-Office or Remote
20-36 Hourly
Junior
20-36 Hourly
Junior
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Provides telephonic after-hours care management for inpatient discharge planning, safe transfers, emergency department diversion, referrals, authorizations, and utilization management. Coordinates with physicians, hospitals, facilities, patients, families, and care teams while documenting interventions, managing caseloads, meeting productivity standards, identifying high-risk patients, and supporting weekend and holiday coverage.
Top Skills: HipaaPc Applications

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account