Palona AI Logo

Palona AI

AI Research Engineer, Computer Vision & VLMs

Posted 2 Days Ago
Be an Early Applicant
In-Office
Los Altos, CA, USA
Mid level
In-Office
Los Altos, CA, USA
Mid level
Develop computer vision and vision-language models for image and video understanding in restaurant environments. Build datasets, training strategies, evaluations, and benchmarks; diagnose model failures; and deploy efficient, reliable inference pipelines. The role combines research and engineering across scene understanding, object tracking, activity recognition, temporal reasoning, visual grounding, and multimodal models, while partnering with product and infrastructure teams to bring research into production.
The summary above was generated by AI

Palona is building AI for the physical world, starting with restaurants. Understanding a busy restaurant means making sense of people, objects, activities, and events as they change over time, despite occlusion, changing lighting, varied camera views, and incomplete information.

We are looking for an AI Research Engineer with a strong research background in computer vision and vision-language models (VLMs) to develop the visual intelligence behind Palona’s products. You will work on image and video understanding, spatiotemporal reasoning, and multimodal models that connect visual observations to useful insights and actions in real restaurant environments.

This role combines research depth with ownership of working systems. You will formulate research questions, build datasets, train and evaluate models, and partner with product and engineering to bring successful approaches into production. Researchers and engineers from autonomous driving, robotics, embodied AI, and related perception fields are especially encouraged to apply.

What you’ll own
  • Develop computer vision and VLM approaches for scene understanding, object detection and tracking, activity recognition, and understanding events across video.
  • Adapt, fine-tune, and evaluate vision and vision-language models for visual grounding, temporal reasoning, and structured prediction grounded in observable evidence.
  • Design training and adaptation strategies, including supervised fine-tuning, representation learning, distillation, and domain adaptation, based on measurable product needs.
  • Build representative image and video datasets, annotation workflows, and evaluation sets that capture difficult edge cases while protecting sensitive data.
  • Create rigorous experiments and benchmarks that measure perception quality, temporal consistency, hallucinations, robustness, latency, and cost across locations and operating conditions.
  • Diagnose failures caused by occlusion, lighting changes, camera placement, rare events, and domain shift; use those findings to improve data and models.
  • Partner with infrastructure and product engineers to deploy efficient inference pipelines, with monitoring, quality gates, staged rollouts, and rollback paths.
  • Translate advances in computer vision, VLMs, and embodied AI into practical product capabilities, and communicate the evidence and tradeoffs behind your decisions.
  • Raise research and engineering standards through reproducible experiments, thoughtful reviews, and clear documentation.

Requirements
  • 3+ years of research or applied development experience in computer vision, multimodal learning, or a closely related field; relevant graduate research counts toward this experience.
  • A demonstrated research track record in computer vision or vision-language modeling, through publications, substantial research projects, open-source contributions, or research delivered in industry.
  • Strong foundations in deep learning, visual representation learning, and experimental design, with depth in areas such as video understanding, detection and tracking, visual grounding, or multimodal reasoning.
  • Hands-on experience training, fine-tuning, or adapting computer vision models, and developing or evaluating VLMs beyond basic API integration.
  • Strong Python skills and experience with PyTorch or an equivalent deep learning framework, along with modern training and evaluation tooling.
  • Experience building datasets, designing reliable evaluations, analyzing model failures, and using ablations to understand what drives improvements.
  • Strong software engineering judgment and the ability to turn research code into reproducible, tested systems that other engineers can use.
  • Ability to connect modeling choices to product constraints including latency, cost, privacy, reliability, and user experience.
  • Comfort working through ambiguity and collaborating across research, engineering, and product.
Especially relevant experience
  • A PhD or research-focused master’s degree in computer vision, machine learning, robotics, or a related field, or equivalent research experience.
  • Industry research or engineering experience in autonomous driving, robotics, embodied AI, or other applications of perception in the physical world.
  • Publications at venues such as CVPR, ICCV, ECCV, NeurIPS, ICLR, ICML, CoRL, ICRA, or RSS.
  • Experience with monocular video perception, spatial understanding, long-video reasoning, or learning from limited and noisy labels.
  • Experience shipping vision models under real-time constraints, including model compression, distillation, quantization, or inference optimization.

When applying, please include links to relevant publications, research projects, or code, and briefly describe your own contribution.


Benefits
  • Competitive salary and stock option plan.
  • Company-sponsored green card applications for strong candidates hired into U.S.-based roles, subject to eligibility.
  • Medical, dental, vision, and retirement benefits as applicable.
  • Family leave and short-term and long-term disability benefits as applicable.
  • Paid time off and company holidays.
  • Learning and development support.
HQ

Palona AI Menlo Park, California, USA Office

Menlo Park, CA, United States, 94025

Similar Jobs

Yesterday
Easy Apply
Remote or Hybrid
Easy Apply
78K-101K Annually
Junior
78K-101K Annually
Junior
Artificial Intelligence • Cloud • Computer Vision • Hardware • Internet of Things • Software
Serve as the primary post-implementation contact for top customers, craft joint success plans, run executive business reviews and workshops, mentor teammates, support French-speaking customers, and advise on customizing Samsara’s IoT platform to drive safety, efficiency, and sustainability.
Top Skills: IotSamsara PlatformVehicle TelematicsVideo-Based Safety
Yesterday
Hybrid
19-22 Hourly
Junior
19-22 Hourly
Junior
eCommerce • Fashion • Retail • Sales • Wearables • Design
Provides friendly, efficient cashier service in a luxury retail store. Greets customers, operates POS and mobile systems, processes purchases, returns, repairs, and gift cards, maintains cashwrap accuracy and cleanliness, completes audits, captures customer information, and promotes products and loyalty programs. Assists with cashier training, resolves equipment issues, follows store policies, and supports customers during busy periods. Requires flexible scheduling, physical mobility, and the ability to lift up to 50 pounds occasionally.
Top Skills: IpadLaptopMobile PosPosWalkie-Talkie
Yesterday
Hybrid
173K-197K Annually
Senior level
173K-197K Annually
Senior level
Fintech • Machine Learning • Payments • Software • Financial Services
Lead hands-on software engineering across the full development lifecycle, owning technical design, architecture, and development of fault-tolerant, cross-functional applications. Establish engineering standards, improve developer tools and frameworks, provide technical guidance, and mentor junior and intermediate engineers. The role requires expertise in application or data architecture, cloud platforms, and modern software development practices.
Top Skills: AgileAWSContainersGoGCPJavaMicroservicesAzureObject-Oriented ProgrammingPythonRestful Apis

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account