Epsilon Health Logo

Epsilon Health

Research Scientist - VLM Pretraining

Reposted 21 Days Ago
Be an Early Applicant
In-Office
San Francisco, CA, USA
Senior level
In-Office
San Francisco, CA, USA
Senior level
Lead design and end-to-end pretraining of large-scale vision-language foundation models for radiology. Responsibilities include VLM architecture design, multimodal data and task mixture engineering, large distributed training runs, fine-grained visual grounding, pretraining evaluation, dataset curation, and handing off base checkpoints for post-training and RL work. Publish research and establish production best practices for medical VLM pretraining at scale.
The summary above was generated by AI
About Us

We're tackling one of healthcare's most critical challenges in medical imaging and diagnostics. Our company operates at the intersection of cutting-edge AI and clinical practice, building technology that directly impacts patient outcomes. We've assembled one of the industry's most comprehensive and diverse medical imaging datasets and have a proven product-market fit with a substantial customer pipeline already in place.

 
Role Overview

We're seeking a Research Scientist with deep expertise in large-scale vision-language pretraining to join our ML Research team. You'll be at the forefront of developing state-of-the-art multimodal models for clinical use in radiology settings. This role owns the pretraining stage of our radiology report generation model: VLM architecture design, multimodal data and task mixtures, and the large-scale training runs that build grounded visual understanding across X-rays, CT scans, and MRI. You'll work with one of the largest and most diverse medical imaging datasets in the industry, paired with the reports that make multimodal pretraining at this scale possible, while maintaining the clinical rigor required for healthcare deployment. Post-training and RL are owned by a partner role you'll collaborate with closely.

Key Responsibilities
  • Design, train, and scale vision-language foundation models for radiology applications, owning the pretraining stage end to end.

  • Develop VLM architectures suited to medical imaging, including native and variable resolution handling, high-resolution tiling, connector design, and token budgets for volumetric studies.

  • Build and tune multimodal pretraining mixtures across captioning, VQA, grounding, and retrieval tasks, balancing data sources to avoid regressions in language capability.

  • Develop fine-grained visual grounding during pretraining, enabling models to localize findings within medical images using bounding boxes or segmentation masks.

  • Own pretraining evaluation (zero- and few-shot transfer, probing, and downstream fine-tunability) — as the signal for base model quality.

  • Train joint vision-language embedding spaces using contrastive and generative objectives, including region- and sentence-level alignment between images and reports.

  • Contribute hands-on to all stages of pretraining including dataset curation, architecture design, distributed training, and handoff of base checkpoints to post-training.

  • Stay current with cutting-edge research in vision-language modeling and large-scale multimodal pretraining.

  • Drive research and technical excellence through conference publications and technical blog posts, establishing best practices for pretraining medical VLMs at scale.

Qualifications
  • 6+ years of academia/industry experience in vision-language modeling, multimodal learning, or related fields

  • Deep expertise in pretraining large vision-language models (e.g., LLaVA, Flamingo, CogVLM, Qwen-VL, InternVL, or similar architectures)

  • Strong foundation in modern VLM pretraining techniques including:

    • Vision-language connector and fusion architectures (projection, cross-attention, resampler-based)

    • Variable and high-resolution image handling (native resolution, dynamic tiling, token compression)

    • Contrastive and generative objectives for learning joint vision-language embedding spaces

    • Data and task mixture design, including curriculum and mixture-ratio ablations

  • Experience with fine-grained visual grounding (referring expression comprehension, phrase grounding, box or mask prediction)

  • Track record of implementing complex models from research papers and adapting them to new domains

  • Proficiency in PyTorch or JAX, with experience training large models on multi-GPU/distributed systems

  • Experience with autoregressive language modeling and long-context training

  • Hands-on experience with medical imaging applications, particularly radiology report generation

  • Strong software engineering skills and ability to write production-quality code

Preferred Qualifications
  • Publications at top-tier conferences (NeurIPS, ICML, ICLR, CVPR, ACL, EMNLP, MICCAI)

  • Experience training vision encoders from scratch, or co-designing them with a downstream VLM

  • Experience with interleaved image-text pretraining and synthetic recaptioning pipelines

  • Experience with 3D medical image processing and temporal modeling

  • Familiarity with clinical NLP and medical knowledge representation

  • Knowledge of evaluation methodologies for long-form generation, including factuality assessment and hallucination detection

  • Experience with model interpretability, explainability, and uncertainty quantification in safety-critical applications

Similar Jobs

12 Minutes Ago
Remote or Hybrid
Sunnyvale, CA, USA
155K-240K Annually
Expert/Leader
155K-240K Annually
Expert/Leader
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Leads CrowdStrike’s global Partner GTM Programs team, developing and optimizing partner programs, incentives, enablement, and emerging go-to-market strategies. Drives revenue, market share, retention, cross-functional alignment, performance analysis, and joint marketing initiatives. Manages a team, advises senior stakeholders, represents the company at industry events, and continuously improves partnership operations. The role requires 10+ years of relevant partner GTM experience and up to 30% domestic and international travel.
Top Skills: AICybersecurity
12 Minutes Ago
Remote or Hybrid
USA
100K-145K Annually
Junior
100K-145K Annually
Junior
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Develop and maintain SOAR playbooks, PowerShell and Python scripts, SIEM integrations, and AI-powered workflows for CrowdStrike’s Falcon Complete MDR operations. The role automates security enrichment, triage, investigation, remediation, and basic forensic tasks while collaborating with SOC analysts and engineering teams. Responsibilities include data parsing, Git-based version control, identifying automation opportunities, and evaluating emerging SOAR and AI technologies to improve analyst efficiency.
Top Skills: Ai Workflow FrameworksAPIsAWSAzureBitbucketCrowdstrike FalconFalcon SoarGCPGenerative AiGitGitGitlabJSONLlmsLogscaleMitre Att&CkNistPowershellPythonRegular ExpressionsSIEMSoar
20 Minutes Ago
Remote or Hybrid
United States
Entry level
Entry level
AdTech • Consumer Web • Digital Media • eCommerce • Marketing Tech • SEO
Supports Medicare business operations across agent licensing and contracting, compliance and call-quality audits, customer retention, member support, enrollment verification, and seasonal initiatives. Maintains licenses, certifications, carrier access, CRM records, and compliance documentation; conducts welcome, remediation, and retention calls; investigates grievances and resolves member issues. The role is remote and requires an active health insurance license, Medicare certification, strong communication, attention to detail, and familiarity with Medicare regulations and carrier systems.
Top Skills: Carrier PortalsCrm PlatformsMulti-State Licensing SystemsQuotemanageSpreadsheetsZendesk

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account