Lead high-impact applied ML research for construction-focused vision-language models. Responsibilities include architecting and training VLMs, developing post-training methods, building video and temporal-reasoning benchmarks, creating scalable training and evaluation pipelines, and optimizing inference through distillation, quantization, and model routing. The role addresses long-video understanding, expert-level perception, efficient large-scale training, and deployment across thousands of hours of daily footage.
About Ironsite
The Role
Open Problems You Could Own in Your First Year
What You'll Do
Technical Challenges You'll Solve
What We're Looking For
What Success Looks Like
Location, Compensation, & Perks
Why You'll Love Working at Ironsite
Ironsite is building the intelligence layer for the physical world. We design our own wearable hardware, deploy it alongside craft workers, and transform a shift's footage into a next-morning report. Our internal team and purpose-built models label the data overnight and deliver actionable insights to superintendents by 5 AM.
We are accelerating the speed, efficiency, and predictability of construction, especially for complex, mission-critical infrastructure projects, including data centers, LNG facilities, sports stadiums, hospitals, and other large-scale developments, by training AI models on egocentric construction footage and labor productivity data. We are built with a pro-worker philosophy at our core: we believe technology should empower the workforce, not replace it. We're working to give craft workers and project leaders better visibility into what's happening on-site, while creating a system where the reality of construction and the chaos of each day is finally available to the people running the project.
Ironsite is deployed across several of the largest active construction projects in the country. To date, we've captured more than 100,000 hours of construction footage across seven states, now process thousands of hours of site activity every day, and maintain a worker opt-out rate below two percent. This is enabled by a workforce-first architecture that anonymizes devices, captures no audio, and never releases raw video.
Ironsite is backed by leading investors (8VC, South Park Commons, Saga Ventures) and prominent operators across technology and construction, including Eric Schmidt, Jeff Dean, Jeff Rothschild, Mark Leslie, Scott Wu, Eric Glyman, Karim Atiyeh, Russell Kaplan, and others, alongside over a dozen construction industry operators who have joined us as partners in building this.
Longer term, we believe Ironsite is the foundation for what construction becomes in the next decade. We think the systems we're building are the operating system for how the physical world gets built, and will unlock a fundamentally different way of respect for our workforce. One where craft workers are more valued, more visible, and better paid for the skill they bring, and where the industry finally has the intelligence layer that makes autonomous construction possible. Both futures start with the same foundation.
We're hiring an exceptional Staff Applied ML Researcher to help build the vision-language models that turn Ironsite's data into the intelligence layer we're building.
You'll work directly with our Chief Science Officer on the research problems that decide how good Ironsite's models become. Training, benchmarking, and deploying state-of-the-art VLMs that can interpret the complexity of a real construction site, built on data that no other lab in the world has access to. This is a role for someone who wants to do frontier research on frontier data, and see their work ship to jobsites where it actually changes how things get built.
Ironsite operates one of the most distinctive research environments in AI today. Our dataset is proprietary, growing by thousands of hours per day, expert-labeled, and structured around a taxonomy built for a specific real-world domain. Our compute footprint spans edge, on-prem, and cloud. And our models don't just ship to a benchmark, they ship to production, running on active jobsites within days of training. Very few research seats in the world offer that combination.
- Long-video understanding at industrial scale. Reasoning over multi-hour egocentric footage where the events that matter are sparse, temporally distant, and often interdependent. This requires temporal grounding, long-context modeling, memory architectures, and evaluation frameworks that go beyond what current VLMs offer. You would own the research direction and set the bar for what's possible.
- Post-training frontiers for expert-level perception. Using SFT, RL (GRPO and beyond), and novel post-training recipes to push VLM performance on fine-grained construction activity to match or exceed expert human taggers across every trade. Not just closing the gap on standard benchmarks, but defining new ones the field will follow.
- Inference at industrial fleet scale. In the next twelve months we will collect and process millions of hours of video per month, and eventually hundreds of thousands of hours per day. Making frontier-quality inference cheap enough to run on all of it, through distillation, quantization, model routing, and hardware-aware optimization, is a partially unsolved problem you would own.
- Evaluation frameworks that predict reality. Building benchmarks that track real field performance and hold the whole team accountable to what those metrics actually predict. Standard academic benchmarks don't cut it for the problems we're solving; you would lead the definition of what does.
- Foundation models for the physical world. No one has trained a foundation model on 5M+ hours of egocentric construction video, because that data has never existed. You would be one of the technical leaders defining how that gets done.
Lead research on Ironsite's core models.
- Design, train, and iterate on vision-language models fine-tuned for spatial intelligence in construction environments. Set the bar for model quality across the research team.
- Own the hardest research initiatives on our roadmap end to end, from problem definition through architecture, training, evaluation, and deployment.
- Push the frontier of what's possible on our data. Establish new baselines, develop novel post-training recipes, invent long-context architectures, and set the technical direction others follow.
Shape the research roadmap.
- Partner directly with the Chief Science Officer to define the research priorities that matter most for Ironsite's next twelve months and next three years.
- Represent research in company strategy discussions, and translate the state of the research into concrete guidance on product, hardware, and business decisions.
- Set the technical bar for the whole research team through the standard of your own work, your reviews of others' work, and your published perspective on what good research looks like at Ironsite.
Own the Construction Intelligence Benchmark suite.
- Lead the design and expansion of our benchmark suite across video question answering, temporal reasoning, activity recognition, and site-level analytical reasoning.
- Define the evaluation frameworks that measure real-world construction task performance, not just standard academic metrics.
- Make Ironsite's benchmark suite the reference point that future construction AI research measures against.
Ship at production scale.
- Apply distillation, quantization, and model routing so state-of-the-art understanding runs affordably across millions of hours of monthly footage.
- Partner deeply with our hardware and infrastructure teams so research decisions and deployment realities inform each other from day one.
- Own the end-to-end story from research direction to production model running on active jobsites, and hold yourself accountable for both.
Grow the research team.
- Mentor Staff and Senior researchers on the team, and help set the culture of intellectual rigor and ambition that defines Ironsite research.
- Partner with talent on identifying and closing the next generation of research hires.
- Contribute to the intellectual culture of the team through paper reading, technical writing, and the standard of thinking you bring to every discussion.
- Training frontier models efficiently under real compute budgets while maximizing performance on the problems that matter for our customers.
- Inventing new pre-training and post-training objectives that capture construction-specific knowledge, temporal reasoning, and fine-grained perception at a level that generic approaches can't reach.
- Designing efficient attention mechanisms and architectural innovations for long-context understanding of construction workflows that span hours, days, and multi-project trajectories.
- Building evaluation frameworks that measure real-world construction task performance beyond standard benchmarks, and defining what "good" means for a category the field hasn't yet formalized.
- Balancing model capability with deployment constraints for edge, on-prem, and cloud inference across a fleet growing by orders of magnitude.
- Working at the seams between research and production. The most interesting research problems at Ironsite live where a training decision propagates all the way through to hardware constraints on a jobsite in Texas. You would be the person who sees the full path and makes the calls that span it.
What We're Looking For
Required
- 8+ years of hands-on experience designing and training large-scale deep learning models, particularly transformer-based architectures, with meaningful time spent operating at a Principal or Staff research level at a company or lab you're proud of.
- A track record of research contributions that meaningfully advanced the state of the art, whether through papers, models, systems, or products that others in the field have built on.
- Deep expertise with modern deep learning frameworks (PyTorch, JAX, or similar) and strong proficiency in Python with solid software engineering fundamentals.
- Deep experience working with and creating large-scale vision or language datasets.
- Comfort operating at the frontier of what's known. You've led research on problems where the right answer wasn't in a paper yet, and you've figured it out anyway.
- Experience mentoring senior researchers and setting technical direction for a team.
- A background in Computer Science, Machine Learning, AI, Robotics, or a related field.
Strongly preferred
- Deep expertise in fine-tuning and post-training large language or vision-language models (SFT, GRPO and other RL methods, parameter-efficient tuning such as LoRA, and novel post-training recipes).
- Hands-on experience with the hardest challenges of video data, including temporal reasoning, long-context modeling, memory, and efficient processing at scale.
- Experience optimizing inference at industrial scale, including quantization, distillation, sparsity, and efficient serving on edge and cloud infrastructure.
- Strong publication record at top-tier AI, ML, or CV conferences, or comparable evidence of research impact (open source, product influence, community leadership).
- Prior experience at a frontier lab, applied AI startup at scale, or research-heavy product company where you shipped research to production.
Nice to have
- Familiarity with MLOps tools for scalable model training and deployment.
- Experience with multimodal models spanning vision, language, and additional modalities (audio, sensor, motion).
- Strong interest in vision-language models applied to real-world physical problems, and genuine curiosity about the day-to-day lives of construction workers.
- First 30 days. You know our data, our benchmarks, and our production models cold. You've identified the highest-leverage research directions for the next twelve months and have made your case for what Ironsite should invest in and why.
- First 3 months. You've led at least one major research initiative from problem definition through deployment. Your work has visibly moved the state of Ironsite's model performance on a problem that matters, and you're actively shaping the research roadmap alongside the Chief Science Officer.
- First 6 months. You're one of the defining research voices at the company. You've mentored the rest of the research team, hired at least one senior researcher, and led the technical direction on a piece of Ironsite's research strategy that will run for years.
- San Francisco Bay Area (on-site)
- Base salary: $250k-$400k per year, commensurate with experience
- Significant early-stage equity
- Full benefits including health, dental, vision, and 401(k) with 6% match
- Access to dedicated GPU compute resources for research and experimentation
- Daily catered breakfast and lunch
- Office in San Francisco, next to Oracle Park and the Caltrain
Final compensation is determined by experience, location, and level.
- Foundational impact. Solve fundamental AI problems to transform one of the world's largest and least-digitized industries. Your models ship to real jobsites, not just papers.
- Ownership and autonomy. As a Principal, you'll have real authority over research direction, technical decisions, and how the team operates. We value intellectual curiosity, first-principles thinking, and iterating quickly to turn ambitious ideas into reality.
- Dream dataset. Exclusive access to a massive, proprietary, and continuously growing corpus of egocentric jobsite video from hundreds of devices deployed on active construction sites. A moat that enables frontier research no other lab can do.
- World-class team. Lead alongside a small, elite team of researchers and engineers who have shipped cutting-edge AI products at companies like DeepMind, Etched, Meta, Apple, and NVIDIA.
The base pay range for this role is $250,000 – $400,000 per year.
Similar Jobs
Artificial Intelligence • Cloud • Information Technology • Consulting
Lead original, publication-driven research in foundational AI, focusing on representation learning, reasoning, physics-inspired ML, and AI-quantum/physical science intersections. Own projects end-to-end: problem formulation, math, implementation, experiments, analysis, and publishing; develop reusable research assets and collaborate across disciplines and external partners.
Top Skills:
C++Density Functional TheoryDistributed SystemsGitHpcMolecular DynamicsMonte CarloPythonPyTorch
Healthtech • Biotech • Pharmaceutical • Manufacturing
Lead clinical research strategy and execution for ophthalmic and vision device programs. Serve as scientific lead, design studies, author clinical documentation, collaborate cross-functionally, engage KOLs and regulators, and manage teams to deliver compliant evidence supporting product development and lifecycle.
Software
Lead state-of-the-art research in long-horizon robot autonomy (TAMP, VLMs, world models, 3D perception, multimodal fusion). Manage researchers and engineers to build and evaluate prototypes on simulated and real robots, collaborate cross-functionally and with external partners, publish papers, and generate patents while shaping R&D priorities.
Top Skills:
C++CudaPyTorchRosTensorFlow
What you need to know about the San Francisco Tech Scene
San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.
Key Facts About San Francisco Tech
- Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
- Major Tech Employers: Google, Apple, Salesforce, Meta
- Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
- Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
- Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
- Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine



