XDOF Logo

XDOF

Senior Perception Engineer

Posted One Month Ago
Be an Early Applicant
Hybrid
San Mateo, CA, USA
Senior level
Hybrid
San Mateo, CA, USA
Senior level
Develop robot perception systems for human pose estimation, 3D vision, sensor calibration, visual SLAM, multimodal synchronization, and 6DoF pose estimation. Train and optimize perception models for embedded deployment using TensorRT and CUDA, build automated annotation and quality-assurance pipelines, and connect perception accuracy with robot training outcomes. Collaborate with robotics, machine learning, and data infrastructure teams to deliver high-quality annotations for teleoperation and embodied AI systems.
The summary above was generated by AI

At XDOF, we’re at an inflection point. Frontier labs are racing to build general-purpose robots, and high-quality training data is the bottleneck. We’re building the foundation behind the foundation models – the data collection systems, operational capability, exabyte-scale data warehouse, and software toolchain – to help our partners drive the field forward.

The Perception Algorithm team transforms raw multimodal sensor data into high-quality robot training annotations. You will be deeply involved in the complete loop from data collection to model delivery — sensor calibration, SLAM localization, human pose estimation, perception model training, and embedded deployment. Your work directly determines the quality ceiling of our training data.

Core Responsibilities

Human Pose Estimation

  • Design and optimize hand pose estimation pipelines supporting accurate joint angle extraction from teleoperation data collection

  • Build full-body pose estimation systems for motion capture and teleoperation action annotation ground truth generation

  • Research and apply vision-based pose estimation methods (markerless) to reduce data collection costs

  • Fuse pose estimation outputs with robot joint angle data to generate consistent training annotations

Robot Perception & Calibration

  • Design and maintain intrinsic/extrinsic calibration pipelines for multi-camera arrays (factory calibration + online recalibration)

  • Build visual SLAM / V-SLAM systems supporting real-time localization and scene reconstruction on data collection platforms

  • Implement hand-eye calibration between cameras and robot end-effectors

  • Develop temporal alignment solutions across multimodal sensors (cameras, IMU, data gloves, force sensors)

Perception Model Training & Deployment

  • Train and iterate on perception models including object detection, instance segmentation, and 6DoF pose estimation

  • Optimize model inference using TensorRT / CUDA for real-time performance on robot embedded platforms

  • Write custom CUDA kernels for low-level acceleration of perception tasks

  • Design evaluation metric frameworks for perception models; continuously track the relationship between model performance and data quality

End-to-End Loop from Data Collection to Model Delivery

  • Contribute to the design of automated annotation pipelines that convert sensor data into structured training labels

  • Build Auto QA modules to filter low-quality data including anomalous frames, failed demonstrations, and sensor dropouts

  • Collaborate with ML engineers and data infrastructure teams to ensure perception output formats meet downstream VLA model training requirements

  • Establish feedback mechanisms linking perception accuracy to model training outcomes, continuously improving annotation quality

Requirements

Must-Have

  • 5+ years of industry experience in robot perception or computer vision

  • Strong 3D vision fundamentals: stereo and structured-light camera principles, 3D reconstruction

  • Proficiency with SLAM frameworks (ORB-SLAM, VINS-Mono, FastLIO, etc.) or V-SLAM system development experience

  • Hands-on engineering experience with human pose estimation: hand joints (MediaPipe, MANO) or full-body pose (OpenPose, SMPLify, etc.)

  • Proficient in deep learning training frameworks for perception model training, tuning, and evaluation

  • TensorRT deployment experience with real-time inference optimization on embedded platforms (Jetson, Horizon, etc.)

  • CUDA programming fundamentals; ability to write or debug custom kernels

  • Proficient in C++ and Python with ROS / ROS2 development experience

  • Proficient with AI coding agents

Nice to Have

  • Engineering experience with 6DoF object pose estimation (FoundPose, FoundationPose, GDR-Net, etc.)

  • Familiarity with 3D Gaussian Splatting or NeRF for scene reconstruction or data augmentation

  • Experience with robot manipulation or teleoperation systems

  • End-to-end development experience with automated annotation pipelines or ground truth generation systems

  • Published research in perception, pose estimation, or robotics

What We Offer

  • Direct involvement in the most critical technical challenge in embodied intelligence: producing high-quality robot training data

  • An environment working alongside top-tier robotics engineers and ML researchers

  • Proprietary hardware platforms (humanoid robots, camera arrays, data gloves)

  • A fast-paced, high-autonomy 0→1 work environment

Similar Jobs

Yesterday
In-Office
150K-200K Annually
Senior level
150K-200K Annually
Senior level
Artificial Intelligence • Hardware • Productivity • Robotics • Software • Automation • Manufacturing
Develop production robotics perception software for manufacturing environments. Responsibilities include 3D geometry reconstruction, segmentation, inspection, active perception, viewpoint planning, multi-sensor fusion, GPU acceleration, hardware-software integration, debugging, and deployment support. The role works directly with physical robots and sensors, assists application teams with proofs of concept and deployments, and travels to customer sites for system-level troubleshooting.
Top Skills: 3D GeometryActive PerceptionC++Computer VisionDeep LearningDefect DetectionGpu ProgrammingIcpImage SegmentationInstance SegmentationMachine LearningMesh ReconstructionMulti-Sensor FusionPoint Cloud RegistrationPoint Cloud SegmentationPoint CloudsPythonRgb-D SensingRobotic SystemsSemantic SegmentationSensor CalibrationSensor Synchronization
2 Days Ago
In-Office
San Jose, CA, USA
Senior level
Senior level
Artificial Intelligence • Automotive • Software • Transportation
Develop and integrate deep-learning computer vision and perception algorithms for autonomous vehicles. Responsibilities include managing data pipelines, designing and optimizing neural networks, defining performance metrics, leading long-term projects, communicating results, and aligning algorithms with product needs. The role requires strong computer vision, deep learning, mathematics, Python, and PyTorch expertise, with C++, Linux, Git, and research publications advantageous.
Top Skills: C++GitLinuxPythonPyTorch
5 Days Ago
In-Office
Santa Clara, CA, USA
152K-288K Annually
Senior level
152K-288K Annually
Senior level
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Develop application software architecture for autonomous-vehicle perception and sensor-fusion systems. Integrate and tune NVIDIA solutions in target vehicles, lead bring-up activities, troubleshoot system issues, analyze in-vehicle and simulation data, and optimize hardware-software performance. Collaborate with architecture, development, partner, and global engineering teams to deploy scalable autonomous-driving solutions.
Top Skills: AspiceAutonomous DrivingCC++CamerasCudaDeep LearningIso 21448Iso 26262LidarNvidia DriveNvidia GpuPerception SystemsPythonQnxRadarSensor FusionUltrasonics

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account