Ando (ando.so) Logo

Ando (ando.so)

Research Engineer

Reposted 28 Days Ago
Be an Early Applicant
In-Office
San Francisco, CA, USA
Senior level
In-Office
San Francisco, CA, USA
Senior level
Founding researcher to develop and evaluate memory architectures and continuous-learning systems for enterprise messaging. Run experiments on offline workspace data, design systems, ship features, and publish research; work balances theory, applied evaluation, and product-focused execution.
The summary above was generated by AI

Ando is a messaging platform where AI agents take on work alongside their human teammates. We’re rebuilding Slack from the ground up around two core ideas: durable memory and agents as first-class participants.

We have real, longitudinal, multi-party workspace data, and human / agent users whose behavior tells you whether the systems you're designing are actually meaningfully improving. Working with live business communication means permissions, redaction, and security are things we have to think about as well. If you want to have your research come into contact with reality, Ando is the place for it.

Some concrete research problems on our plate right now:

  • Proactivity: An agent embedded in a team's channels has to decide, message by message, whether to ignore, quietly track, or intervene. Evaluating that judgment means building benchmarks where the ground truth includes silence. Existing agent benchmarks are almost entirely reactive.

  • Memory: What should a workspace agent remember across weeks and months of participation, in what representation, and how do you measure whether memory is helping versus hurting agent usefulness?

  • Continual Learning: How much does basic memory, retrieval, and context impact agent performance, vs where do we genuinely need RL and continual learning?

What you’ll be doing at Ando

You'll be one of the first members of a small research team, working close to both the data and the product.

  • Build evaluation from real data - Mine production workspace data (carefully, with consent and redaction pipelines you'll help design) into benchmarks and labeled datasets. Design label schemas, run labeling with real inter-rater rigor, and build the harnesses that make expert judgment cheap to capture and hard to corrupt.

  • Run experiments that ship - Initial work happens on offline workspace data; the destination is production systems used by every Ando customer. The distance between "the benchmark improved" and "the feature shipped" should be weeks and you'll own both ends.

  • Decide what we build versus who we partner with - We won't do everything in-house. Part of the job is evaluating frontier vendors and research teams (eval infrastructure, observability, continual-learning tooling) and choosing who we build with.

  • Publish when we have something real - We expect the team to publish as we make meaningful progress. But ideas are not Ando's moat; execution, product quality, and customer experience are. Research here is in service of customers first, and the publications will be better for it: they'll describe things that actually worked on real data.

What we're looking for
  • Strong applied research background, with depth in model evaluation, benchmarking, and/or failure analysis. You've built evals you trusted enough to make decisions with.

  • Evidence over credentials. Work samples or code that demonstrate the skills: eval frameworks, benchmark suites, failure-analysis reports or tooling, labeling infrastructure. Show us something you built to find out whether a system actually worked.

  • Strong technical communication. You can explain complex ideas simply and hold high-bandwidth, generative technical conversations with researchers and with our product team.

  • Comfort with mess. Real workspace data is incomplete, ambiguous, and full of edge cases that break clean abstractions. You treat that as signal, not noise.

  • Bonus: familiarity with simulation (Park et al.), human-in-the-loop evaluation (Scale HIL leaderboard), Cartridges and related context/memory-compression work, memory for multi-party long-running settings, or agent observability standards (setting up Langsmith or similar).

Hiring process

  1. 30 min intro call

  2. Technical conversation - walk us through an app you've shipped

  3. Paid take-home or IRL work trial

Benefits

  • Free Equinox membership & other health perks

  • Generous equity grant vested over 4 years

  • Health, dental, vision insurance

HQ

Ando (ando.so) San Francisco, California, USA Office

300 Grant Ave, San Francisco, California, United States, 94108 3629

Similar Jobs

8 Days Ago
Hybrid
Santa Clara, CA, USA
203K-354K Annually
Senior level
203K-354K Annually
Senior level
Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Lead research in agent learning and recursive self-improvement for enterprise AI agents. Design post-training methods, agent harnesses, training environments, evaluations, and distributed pipelines across language and multimodal systems. Analyze failures, develop training signals, run rigorous experiments, and transition validated improvements into reliable production systems. Collaborate with research, engineering, infrastructure, security, and product teams while communicating results through publications, patents, reports, and open-source work.
Top Skills: DeepspeedFsdpKubernetesMcpMegatronOpenrlhfPythonPyTorchPytorch DistributedRaySglangSlurmTrlVerlVllm
9 Days Ago
Hybrid
Santa Clara, CA, USA
232K-405K Annually
Senior level
232K-405K Annually
Senior level
Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Lead high-impact research in agent learning and recursive self-improvement. Design post-training methods, agent harnesses, training environments, evaluations, and distributed pipelines for reliable enterprise agents. Investigate planning, reasoning, memory, tool use, multimodal capabilities, and long-horizon execution. Analyze failures, run rigorous experiments, build improvement loops, and transition validated research into production systems. Provide technical leadership across teams and communicate results through reports, publications, patents, benchmarks, and open-source contributions.
Top Skills: Distributed InferenceDistributed TrainingDpoGpu InfrastructureGrpoKnowledge RetrievalLarge Language Models (Llms)Model Context Protocol (Mcp)Multimodal ModelsPythonPyTorchReinforcement LearningReward ModelingSupervised Fine-Tuning (Sft)
16 Days Ago
Hybrid
San Francisco, CA, USA
136K-265K Annually
Junior
136K-265K Annually
Junior
Cloud • Healthtech • Social Impact • Software • Biotech
Build datasets, evaluations, and scalable data infrastructure to improve frontier AI models for scientific and biological applications. Analyze model failure modes, run experiments, translate scientific expertise into rigorous evaluation criteria, and collaborate with scientists and AI labs. The role requires working at the intersection of software engineering, biology, and frontier AI in an in-person, fast-paced San Francisco environment.
Top Skills: AIData InfrastructureData PipelinesFrontier ModelsLarge Language Models (Llms)

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account