Apollo Research Logo

Apollo Research

Full-stack Software Engineer (Product)

Posted 6 Days Ago
In-Office
San Francisco, CA, USA
149K-290K Annually
Mid level
In-Office
San Francisco, CA, USA
149K-290K Annually
Mid level
Build and operate Watcher, a full-stack coding-agent security product. Develop backend APIs, ingestion and grading pipelines, agent hooks, CLI tooling, monitoring dashboards, enterprise integrations, and real-time blocking systems. Own features from customer discovery through production deployment, while ensuring reliability, scalability, tenant isolation, encryption, access controls, and flexible cloud, on-premises, and local deployments. Collaborate with customers and security teams to improve the product and productionize monitoring research.
The summary above was generated by AI
THE OPPORTUNITY

We are building Watcher, a coding agent security product. Watcher is deployed in production and monitors billions of agent tokens per month across engineering teams at agent-building scale-ups and enterprises. We are looking for a product engineer to build and scale Watcher end to end.

This role includes backend and frontend components. You will own features across the full stack, from agent hooks and ingestion pipelines to monitoring dashboards and enterprise integrations. You can expect to  own features from customer conversation to production deployment.

This is truly a “start-up role”: you will own big chunks of the product, make decisions fast, and ship at high volume. You will join a small team with significant ability to shape the product and tech, and you can earn more responsibility quickly. This is an individual contributor role but could lead to management responsibilities eventually, if desired.

KEY RESPONSIBILITIES

    Ship product end to end (~45%)

    Feature development 

  • Design, build, and ship Watcher features across the stack: agent-side hooks and CLI tooling, ingestion and grading pipelines, real-time monitoring UI (Watcher Live), and the organization-wide Analyzer dashboard.

  • Own features end to end, from ambiguous requirements to production: talk to users, scope the work, make the technical decisions, put that into production, and iterate based on usage.

  • Build Watcher integrations:  We currently support Claude Code and Codex, but we’re planning on increasing our coverage to other coding agents, LLM gateways, and anywhere that could benefit from agent monitoring. 

  • Build the policy and configuration layer that lets security teams set org-wide guardrails for  engineers to operate safely within.

  • Productionize research

  • Turn monitoring research (new monitors, grading strategies, backtesting results) into production-ready product features. We are "our own customer": You will use the product you build every day.

  • Infrastructure and scale (~35%)
  • Design and operate backend systems that process large volumes of agent logs in real time.

  • Own reliability: robust error handling, graceful degradation, and observability (logging, metrics, tracing, alerting). We should catch issues before customers do.

  • Architect data models and storage that handle both high-throughput writes and complex analytical queries over historical trajectories. Retention policies have to fit highly sensitive customer data.

  • Make sure Watcher is safe to deploy: Work with our AI Security & Control Engineer on tenant isolation, encryption, and access controls. Watcher touches some of the most sensitive data our customers have.

  • Customer and integration engineering (~20%)
  • Build and maintain integrations with the enterprise security stack, e.g. streaming monitoring alerts into SIEM systems and existing security operations workflows.

  • Support flexible deployment models: cloud-hosted, on-prem, and local backends, so sensitive agent logs never have to leave customer control.

  • Talk to customers, debug issues in their environments, and feed what you learn back into the roadmap.

REPRESENTATIVE PROJECTS

  • Secure a new frontier coding agent: Take an agent we don't yet support (e.g. Cursor) from zero to fully monitored. Build the hook/integration layer, normalize its trajectory format into our data model, validate monitor accuracy on real traffic, and ship it to customers.

  • Build the real-time blocking path: Design the low-latency pipeline that lets monitors block dangerous actions (e.g. git push --force, secret exfiltration) before they execute, while keeping p95 latency low enough that developers don't feel it. Handle traffic spikes and partial failures gracefully.

  • Ship an Analyzer capability end to end: Design and build an organization-wide view (e.g. failure trends over time, cross-session pattern detection) from data model through API to UI, based on what security teams actually need.

  • Deploy Watcher inside locked-down enterprises: Build and document the self-hosted/on-prem deployment story so that enterprises with strict data requirements can run Watcher  inside their own infrastructure.

JOB REQUIREMENTS

    Must-haves
  • 4+ years building production software. You have shipped and operated real systems with real users, and you know what it takes to keep them reliable.

  • Full-stack capability. You are strong on the backend (APIs, data pipelines, databases, cloud infrastructure) and productive on the frontend (modern web UI). You don't need to be world-class at both, but you must be able to own a feature across the whole stack. Our stack is primarily Python and TypeScript.

  • High agency and ownership. You are willing to own big chunks of the product, make decisions fast with incomplete information, and be accountable for the outcome. You don't wait to be told what to do and have an accurate sense of the roadmap for the product.

  • High output volume. You have used Claude Code, Codex, Cursor, or similar tools heavily to accelerate or have built agents or agent tooling yourself.

  • Startup pace. You are excited about a fast-moving environment, comfortable with ambiguity and changing priorities, and willing to grind when it matters.

  • Product sense. You can talk to users, translate vague requirements into concrete designs, and make good calls about what to build and what to cut.

  • Strong nice-to-haves
  • Previous work on developer tools, monitoring/observability systems, or security products. Watcher sits at the intersection of all three.

  • Real-time or large-scale data processing experience (streaming pipelines, message queues, high-throughput log systems).

  • Enterprise or on-prem deployment experience. Shipping software into environments you don't control is its own skill.

  • Experience building LLM-powered applications, e.g. LLM-as-judge setups, evaluation pipelines, or familiarity with frameworks like Inspect.

  • Early-stage startup experience. You have been employee 1-15 somewhere and know what it feels like to build a product with no playbook.

  • Explicitly not required
  • Formal AI safety background. We need excellent product engineers who can learn the AI safety context, not AI safety researchers who need to learn engineering.

  • Management experience. This is an IC role, at least initially.

  • Deep experience in every technology we use. We care about demonstrated ability to learn and ship, not checkbox familiarity with our exact stack.

BENEFITS

  • This role offers market competitive salary, equity, and competitive benefits.

  • Salary: San Francisco: $222,000 – $290,000 London: £149,000 – £195,000. We will be looking to meaningfully raise salaries soon.   

  • Our engineers effectively have an unlimited token budget. If a better result costs more compute, use it.

  • Flexible work hours and schedule

  • Unlimited vacation

  • Unlimited sick leave

  • Up to 6 months of paid parental leave

  • Comprehensive health, dental and vision insurance

  • Retirement savings with competitive employer matching (e.g. 401(k) for US employees)

  • Lunch, dinner, and snacks are provided for all employees on workdays

  • Paid work trips, including staff retreats, business trips, and relevant conferences

  • A yearly $1,000 (USD) professional development budget

  • Relocation support and visa fees (if applicable)

LOGISTICS

  • Time Allocation: Full-time

  • Location: This is an in-person role working out of our London or San Francisco office. We offer flexible working hours and some wfh arrangements.

  • Visa sponsorship: We sponsor visas in both the UK and US. Sponsorship isn't guaranteed for every role or candidate, but if we make you an offer, we'll work with you to find the right visa route.

ABOUT THE TEAM

The product team consists of product engineers: Jeremy Neiman, Zak Walters, Zen van Riel, Srdjan Miletic and Gustavo Bicalho; research scientists: Victor Gillioz, Monika Jotautaitė, Dmitrii Volkov; and our GTM lead: Kyle Dai. Marius Hobbhahn (CEO) advises the team. Furthermore you will interact with our other SWEs and researchers, since we intend to be "our own customer" by using our products internally for our research work. You can find our full team here.

ABOUT APOLLO RESEARCH

The rapid rise in AI capabilities offers tremendous opportunities, but also presents significant risks. At Apollo Research, we're primarily concerned with risks from Loss of Control, i.e. risks coming from the model itself rather than e.g. humans misusing the AI. We're particularly concerned with deceptive alignment / scheming, a phenomenon where a model appears to be aligned but is, in fact, misaligned and capable of evading human oversight. 

We work on the science of scheming, detection of scheming (e.g. building evaluations), and scheming mitigations (e.g. anti-scheming). We also work on control and monitoring research (see our scalable monitoring agenda). We work closely with many frontier AI companies, such as OpenAI, Anthropic, Google, Meta, Thinking Machines and others, e.g. to test their models and collaborate on the science of scheming. At Apollo, we aim for a culture that emphasizes truth-seeking, being goal-oriented, giving and receiving constructive feedback, and being friendly and helpful. If you're interested in more details about what it's like working at Apollo, you can find more information here.

We also build a coding agent security product called Watcher that secures agent deployments in companies. Our goal is to reduce the probability of catastrophic incidents by securing coding agents, learning about their real-world risks, and publishing our research on how to build these control systems most effectively.

Equality Statement: Apollo Research is an Equal Opportunity Employer. We value diversity and are committed to providing equal opportunities to all, regardless of age, disability, gender reassignment, marriage and civil partnership, pregnancy and maternity, race, religion or belief, sex, or sexual orientation.

HOW TO APPLY

Please complete the application form with your CV. The provision of a cover letter is neither required nor encouraged. Please also feel free to share links to relevant work samples.

About the interview process: Our multi-stage process includes a screening interview, a take-home test (3 hours), 3 technical interviews, and a final interview with Marius (CEO). There are no leetcode-style general coding interviews. You may use AI tools on the take-home; we judge the result the way we'd judge any contributor's work, so you are responsible for the quality of everything you submit. If you want to prepare, we suggest building simple monitors for coding agents and running them on your own Claude Code / Cursor / Codex / etc. traffic.

Your Privacy and Fairness in Our Recruitment Process: We are committed to protecting your data, ensuring fairness, and adhering to workplace fairness principles in our recruitment process. To enhance hiring efficiency, we use AI-powered tools to assist with tasks such as resume screening. These tools are designed and deployed in compliance with internationally recognized AI governance frameworks. Your personal data is handled securely and transparently. All resumes are screened by a human and final hiring decisions are made by our team. If you have questions about how your data is processed or wish to report concerns about fairness, please contact us at [email protected].

Similar Jobs

6 Days Ago
In-Office
118K-158K Annually
Mid level
118K-158K Annually
Mid level
Digital Media • Gaming • News + Entertainment • Sports
Develop, test, maintain, and document scalable software features for Disney’s news and entertainment platforms. The role contributes to architecture, code reviews, debugging, technical planning, and distributed systems development. It collaborates across technical teams, supports Agile practices, helps onboard teammates, translates requirements into technical tasks, and serves as an escalation point for technical issues. Core work includes Java and full-stack development, cloud infrastructure, event-driven messaging, databases, monitoring, and automated deployment.
Top Skills: AgileAmazon S3Amazon SnsAmazon SqsAWSAws Ecs FargateAzureDatadog ApmDynamoDBElasticacheElasticsearchGCPGitlab Ci/CdInfrastructure As CodeJavaJavaScriptOpensearchPythonRedisRelational DatabasesScrumSpring BootSpring FrameworkTerraform
17 Days Ago
In-Office
Los Gatos, CA, USA
250K-413K Annually
Senior level
250K-413K Annually
Senior level
News + Entertainment
Design and build scalable, low-latency full-stack applications and UIs for localization workflows. Collaborate with design, product, and engineering to deliver React-based frontends, consume GraphQL/REST APIs, and work across services (Java/OO languages). Drive technical design, performance optimization, and production reliability while integrating applied AI where beneficial.
Top Skills: AICSSGraphQLHTMLJavaJavaScriptLlmsReactRestful ApisTypescript
23 Days Ago
In-Office
San Francisco, CA, USA
Internship
Internship
Artificial Intelligence • Machine Learning • Software
Work on Netic's AI product team to design, code, and ship full‑stack features. Collaborate with customers to gather feedback, own end‑to‑end delivery from data models and APIs to frontend polish, and ensure production reliability through testing and monitoring during a 12‑week internship.
Top Skills: Cloud InfrastructureDatabasesEmbeddingsLlm ApisPythonRagReactTypescript

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account