Middesk Logo

Middesk

Lead Data Scientist

Reposted One Month Ago
Be an Early Applicant
Hybrid
San Francisco, CA, USA
210K-250K Annually
Senior level
Hybrid
San Francisco, CA, USA
210K-250K Annually
Senior level
Responsible for building ML applications focused on risk and fraud, tackling data challenges, innovating in feature engineering, and establishing ML infrastructure.
The summary above was generated by AI
About Middesk:

Middesk is building the data and intelligence infrastructure that helps businesses work together with confidence. We started by creating a comprehensive platform for understanding businesses, bringing together authoritative and proprietary data to help customers verify business identities, onboard customers faster, and manage risk throughout the customer lifecycle.

Today, Middesk is used by more than 700 banks and fintechs, and in 2025 we verified more than 7 million companies. We've also expanded beyond business verification to help companies form, register, manage, and maintain their businesses, supporting more than 50,000 companies in setting up over 100,000 accounts required to hire employees, run payroll, and stay compliant.

Middesk came out of Y Combinator, and is backed by Sequoia Capital, Accel, Insight Partners, and Canapi. We're proud to be named on the Forbes Fintech 50 and Best Startup Employers lists.

About The Role:

We are actively building AI-driven applications that streamline customer workflows, focusing on business onboarding. With our proprietary identity data assets and deep domain expertise, we are uniquely positioned to expand into a broader set of AI-powered solutions that drive long-term growth.

We’re looking for a hands-on applied ML expert to help build the technical foundation for these efforts. Ideally you have shipped external-facing models in the risk/fraud space and know the messy realities of imbalanced data, low labels, and changing behavior. This is a highly technical, hands-on role with wide influence on how we design, build, and scale ML at Middesk.

We follow a hybrid work model, and for this role, there is an expectation of 2 days per week in our SF/NYC office. Candidates should be based within a commutable distance, as we believe in the value of in-person collaboration and building strong team connections while also supporting flexibility where possible.

What You'll Do:
  • Build risk & fraud ML applications: Deliver production ML models in fraud, trust & safety, KYB, and compliance domains, with measurable impact on customer workflows.

  • Tackle hard data problems: Work on classification problems with extreme class imbalance, sparse signals, and “cold start” label challenges.

  • Innovate in feature engineering & labeling: Use graph-based techniques, weak supervision, LLMs, and AI agents to improve signal extraction and automate labeling process.

  • Establish ML infrastructure foundations: Partner with the ML infra team to design feature services, model training pipeline, model serving standards, and orchestration to scale multiple ML use cases.

  • Design and implement knowledge graph solutions: Leveraging LLMs for graph construction, querying, and retrieval to enhance entity resolution and business identity use cases.

What We're Looking For:
  • 7+ years of production ML experience in one or more of the following areas:

    • Building Production ML for risk, fraud, credit, or trust & safety: Track record of shipping external-facing ML applications in one or more of these domains.

    • Knowledge graph applications: Hands-on experience building, querying, or extracting signals from knowledge graphs—ideally over business entity networks (companies, persons, addresses, relationships) to support identity verification, fraud detection, or risk decisioning.

    • Entity resolution for business or individual identities: Experience disambiguating and linking records across noisy, incomplete, or conflicting data sources—particularly in KYB, KYC, AML, or identity verification contexts where the same real-world entity may appear under different names, addresses, or tax IDs.

  • Expertise in classification with real-world ML challenges, for example: imbalanced labels, sparse signals, cold start, and production version management.

  • Hands-on ML infrastructure experience: feature stores, model management, ML training/serving pipelines.

  • Comfort as a senior IC: setting technical direction, mentoring peers, and establishing best practices.

Nice-To Have:
  • B2B SaaS experience, ideally building ML products for enterprise customers.

  • ML pipeline and automation engineering: Experience building end-to-end training harnesses that automate feature engineering, data validation, and model training.

  • Experience scaling ML across multiple products or risk domains.

HQ

Middesk San Francisco, California, USA Office

85 2nd St, Suite 710, , San Francisco, California , United States, 94105

Similar Jobs

Yesterday
Easy Apply
In-Office
Easy Apply
Senior level
Senior level
Artificial Intelligence • Healthtech • Software • Telehealth
Lead patient lifecycle analytics for a digital health product. Build production-ready churn, resurrection, and targeting models; design experiments and apply causal inference; deliver actionable analyses with quantified opportunities; define product metrics; and contribute to the company-wide experimentation platform. Partner closely with Product, Engineering, Design, and Growth teams to shape priorities and improve patient outcomes.
Top Skills: PythonSQL
Yesterday
Hybrid
2 Locations
195K-343K Annually
Senior level
195K-343K Annually
Senior level
Artificial Intelligence • Cloud • Machine Learning • Mobile • Software • Virtual Reality • App development
Leads experimentation, causal inference, advanced statistical analysis, machine learning, and scalable data-product development. Partners with product, engineering, and research teams to generate insights, identify opportunities, influence executive decisions, and shape product strategy. Owns cross-functional projects, analytics roadmaps, reporting, stakeholder adoption, and methodological standards while translating complex data into actionable business outcomes.
Top Skills: A/B TestingAi ToolsCausal InferenceData Orchestration PipelinesEconmlMachine LearningPythonRScipySQLStatsmodels
27 Days Ago
Hybrid
90K-160K Annually
Mid level
90K-160K Annually
Mid level
Big Data • Fintech • Information Technology • Business Intelligence • Financial Services • Cybersecurity • Big Data Analytics
Leads client-facing data science engagements for financial services customers, developing predictive risk management and business intelligence solutions across consumer lending. Performs descriptive, predictive, and prescriptive analysis using advanced statistical techniques and large-scale datasets. Programs data extraction and analysis in R, SAS, SQL, Hive, and Pig; presents recommendations to executives and customers; supports sales proposals and product adoption; mentors junior colleagues; and manages multiple initiatives with limited supervision.
Top Skills: C++HadoopHiveJavaLinuxMS OfficePigPythonRSASSQL

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account