Goaly Logo

Goaly

Founding AI Researcher, RL

Posted 5 Hours Ago
Be an Early Applicant
In-Office
Palo Alto, CA, USA
Entry level
In-Office
Palo Alto, CA, USA
Entry level
Own the research-engineering loop for turning base models into capable, reliable agents. Design RL environments, training and evaluation data, post-training experiments, reward systems, and regression suites. Analyze model behavior, diagnose reward hacking and other failures, improve distributed training and rollout infrastructure, and translate successful research into reproducible production pipelines. The role requires strong Python and software-engineering skills, modern language-model experience, deep-learning knowledge, and reinforcement-learning intuition.
The summary above was generated by AI
About us

We’re building toward a world where every company can become its own AI lab.
Goaly is a stealth AI startup founded by ex-Meta MSL engineers and researchers. Our mission is to dramatically lower the cost, time, and talent barriers to building proprietary AI — and make each generation of models faster and cheaper to build than the last.

Backed by leading AI investors and endorsed by frontier AI researchers and builders, we’re looking for exceptional new grads who want to work on hard, foundational AI systems problems with outsized ownership from day one.

About the role

You will own the experimental loop that turns a capable base model into a useful agent. You will design tasks and environments, prepare training and evaluation data, run reinforcement-learning and related post-training experiments, diagnose model behavior, and convert results into better recipes and production models.

This is a research-engineering role. The best candidates are equally comfortable forming hypotheses, writing high-quality code, operating training pipelines, and investigating why a model or metric moved. You will work closely with RL systems, training, inference, product, and domain experts; when infrastructure slows the science, you will help improve the infrastructure rather than treating it as someone else's problem.

What you'll do
  • Design and run post-training experiments for agentic capabilities, including tool use, coding, reasoning, planning, long-horizon task completion, and recovery from failure.

  • Prepare high-quality training and evaluation data: define task distributions, curate and filter examples, control contamination, balance difficulty, and build reproducible data-generation pipelines.

  • Build realistic RL environments and task harnesses with clear interfaces, reliable resets, isolated execution, useful telemetry, and reward signals that are hard to game.

  • Develop evaluations that measure both capability and reliability. Create regression suites, behavioral slices, error taxonomies, and dashboards that connect aggregate metrics to concrete model failures.

  • Iterate on training recipes, including supervised warm starts, sampling strategies, reward design, verifiers, curricula, optimization choices, and reinforcement fine-tuning methods.

  • Analyze trajectories and model behavior to find reward hacking, shortcut learning, mode collapse, distribution gaps, and other failure modes; turn those findings into targeted experiments.

  • Improve the research workflow through better experiment configuration, rollout inspection, reproducibility, checkpoint evaluation, and automated comparison of runs.

  • Partner with systems engineers to debug cross-layer problems in rollout inference, environment execution, distributed training, and data movement.

  • Translate successful research ideas into stable, repeatable pipelines and help set the team's longer-term post-training roadmap.

You may be a good fit if you have
  • Strong Python and software-engineering skills, including the ability to turn ambiguous research ideas into reliable experimental systems.

  • Hands-on experience training, fine-tuning, or evaluating modern language models, or an exceptional record in a closely related ML research area.

  • Solid understanding of deep learning and optimization, plus enough reinforcement-learning intuition to reason about policies, rewards, sampling, credit assignment, and evaluation bias.

  • Excellent experimental judgment: you define controls, inspect data, validate metrics, keep results reproducible, and distinguish a real improvement from noise or leakage.

  • Ability to debug across model behavior, data, code, and distributed infrastructure without losing sight of the user-facing capability being improved.

  • Clear written and verbal communication and a track record of productive collaboration across research and engineering.

Strong pluses
  • Experience with RLHF, reinforcement fine-tuning, preference optimization, reward or verifier modeling, or large-scale online sampling.

  • Experience building agent environments, secure sandboxes, coding benchmarks, tool-use tasks, or long-horizon evaluations.

  • Familiarity with PyTorch or JAX and distributed ML systems; experience with frameworks such as FSDP, Megatron, DeepSpeed, Ray, VeRL, or related stacks.

  • A record of influential research, open-source contributions, technically ambitious independent projects, or production model launches.

How we work
  • Mission first. We choose work for its impact on the mission and take responsibility for the outcome, not just our assigned tasks.

  • High agency. We identify what is missing, form a plan, and move without waiting for perfect clarity.

  • Speed with rigor. We ship, measure, and iterate quickly while protecting correctness, safety, and reliability.

  • Flexible scope. We cross team and technical boundaries when that is the fastest way to solve the real problem.

  • Low ego, high standards. We give direct feedback, change our minds when the evidence changes, and help the whole team win.

  • Continuous learning. The stack changes quickly; we are willing to learn unfamiliar systems, methods, and domains as the work demands.

Location, visa sponsorship & benefits

  • Hybrid in Palo Alto: 4+ days/week in office.

  • Visa sponsorship: H-1B and OPT/CPT support available, with immigration counsel.

  • Meals & perks: Complimentary lunch, dinner, snacks, and drinks.

A note on qualifications. We value exceptional ability over perfect keyword matches. If the work excites you and you can show strong technical ability, learning speed, or ownership, we encourage you to apply.

Equal opportunity

We are an equal opportunity employer. We consider qualified applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, disability, veteran status, genetic information, or any other characteristic protected by applicable law. We provide reasonable accommodations for candidates who need them during the hiring process.

Similar Jobs

10 Minutes Ago
In-Office
106K-232K Annually
Senior level
106K-232K Annually
Senior level
Aerospace • Information Technology • Software • Cybersecurity • Design • Defense • Manufacturing
Designs, integrates, tests, debugs, and leads development of embedded flight software for satellite and payload systems. Responsibilities include software-hardware integration, requirements analysis, cyber vulnerability analysis, cyber monitoring algorithms, configuration automation, documentation, quality assurance, and delivery coordination. The role interfaces with multidisciplinary engineering teams and supports safety, security, and performance objectives for commercial and government space programs.
Top Skills: BitbucketCC++ConfluenceCyber Monitoring AlgorithmsDevOpsEmbedded SystemsGitlabJavaJIRAPythonReal-Time SoftwareWireshark
11 Minutes Ago
In-Office
177K-239K Annually
Expert/Leader
177K-239K Annually
Expert/Leader
Aerospace • Information Technology • Software • Cybersecurity • Design • Defense • Manufacturing
Leads electrophysics and radar engineering for air and space payloads across the product lifecycle. Responsibilities include radar performance and trade studies, signal and image processing, RWR integration, mission design, payload effectiveness analysis, technical leadership of teams and subcontractors, anomaly resolution, and capability roadmap development. The role requires advanced radar or signal-intelligence expertise, systems integration experience, and active Top Secret/SCI clearance.
Top Skills: Image ProcessingMatlabRadar SystemsRadar Warning Receivers (Rwr)Rf SystemsRfi MitigationSignal ProcessingStkSynthetic Aperture Radar (Sar)
11 Minutes Ago
In-Office
120K-198K Annually
Senior level
120K-198K Annually
Senior level
Aerospace • Information Technology • Software • Cybersecurity • Design • Defense • Manufacturing
Design and validate ASIC and mixed-signal subsystem architectures for space-based digital communication systems. Responsibilities include requirements derivation, verification planning, ADC/DAC performance analysis, technical trade studies, supplier coordination, hardware and software integration, qualification, unit sell-off, and risk mitigation. The role supports the full subsystem lifecycle and requires collaboration across engineering, program management, suppliers, and customers.
Top Skills: AdcAsicCDacI2CJesdJtagMatlabPythonSerdesSpiVlsi

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account