Weekday, Inc. Logo

Weekday, Inc.

LLM Red Team Specialist - Failure Modes & Edge Cases

Reposted 22 Days Ago
Be an Early Applicant
Remote
Hiring Remotely in United States
60-90 Hourly
Junior
Remote
Hiring Remotely in United States
60-90 Hourly
Junior
Design and run red-team evaluations to find failure modes and edge cases in frontier LLMs. Create reproducible, multi-step benchmark tasks, document technical findings, and collaborate with researchers to refine grading and improve model robustness. Commit ~35 hours/week as a remote independent contractor.
The summary above was generated by AI

This role is for one of our clients

Compensation: $60-$90 per hour

Join a pioneering AI initiative focused on building next-generation evaluation benchmarks for frontier AI models. We are seeking analytical and technically skilled professionals to identify where advanced AI systems fail in subtle, real-world scenarios. Working in a red-teaming environment, you will design challenging, multi-step tasks that expose hidden vulnerabilities, reasoning gaps, and edge cases that traditional evaluations often miss.

In this role, you'll collaborate closely with AI researchers to transform discovered failure modes into high-quality benchmark tasks that improve the robustness, safety, and reasoning capabilities of state-of-the-art AI systems.

This is a fully remote, full-time engagement requiring approximately 35 hours per week.


RequirementsKey Responsibilities
  • Investigate how frontier AI models perform across coding, machine learning, analytical reasoning, and complex problem-solving tasks.
  • Identify hidden failure modes, edge cases, reasoning errors, and vulnerabilities that may not be apparent through standard testing.
  • Design challenging evaluation tasks that accurately measure AI capabilities while remaining objective and reproducible.
  • Document findings with clear technical explanations, supporting evidence, and reproducible methodologies.
  • Collaborate with benchmark designers and AI researchers to refine evaluation tasks, eliminate loopholes, and strengthen grading criteria.
  • Share insights and recommendations with cross-functional teams to continuously improve AI evaluation quality and benchmark coverage.
Required Qualifications
  • Master's degree, PhD, or equivalent practical experience in a STEM discipline involving research, coding, or advanced data analysis.
  • Minimum 1 year of experience in AI research, research engineering, security research, AI evaluation, or a related technical field.
  • Demonstrated experience identifying vulnerabilities, adversarial behaviors, edge cases, or failure modes in Large Language Models or other machine learning systems.
  • Strong proficiency in Python and Git, with the ability to build custom scripts for experimentation, testing, and analysis.
  • Solid understanding of modern Large Language Models, their strengths, limitations, and evaluation methodologies.
  • Experience with AI benchmarking, model evaluation, adversarial testing, prompt engineering, or dataset creation is highly desirable.
  • Excellent analytical thinking, creativity, and attention to detail, with the ability to solve ambiguous, open-ended problems independently.
  • Outstanding written communication skills for documenting technical findings clearly and accurately.
  • Ability to commit approximately 35 hours per week on a consistent basis.
Preferred Qualifications
  • Experience with AI safety, red teaming, adversarial machine learning, or security research.
  • Background in benchmark design, evaluation framework development, or AI quality assurance.
  • Experience creating reproducible technical experiments and documenting complex failure analyses.
  • Familiarity with frontier AI research methodologies and model capability assessments.
Why Join
  • Help shape the future of AI evaluation by identifying critical weaknesses before they reach production.
  • Work on cutting-edge AI systems alongside researchers developing next-generation language models.
  • Apply your technical expertise to improve AI reliability, reasoning, and robustness.
  • Contribute directly to benchmark development that influences the evolution of advanced AI technologies.
  • Enjoy the flexibility of a fully remote engagement while working on impactful research initiatives.
Equal Opportunity

We are committed to fostering an inclusive and diverse environment where all qualified applicants receive equal consideration. Reasonable accommodations are available throughout the application and engagement process.

Contract & Engagement Details
  • Independent contractor engagement.
  • Fully remote with flexible working hours.
  • Expected commitment of approximately 35 hours per week.
  • Project duration may be extended, shortened, or concluded based on project requirements and individual performance.
  • Work does not require access to confidential or proprietary information from any current or former employer.
  • Payments are issued weekly based on approved work completed.
  • At this time, we are unable to support H1-B or STEM OPT candidates.

Similar Jobs

22 Minutes Ago
Remote or Hybrid
San Jose, CA, USA
87K-147K Annually
Entry level
87K-147K Annually
Entry level
Aerospace • Hardware • Information Technology • Security • Software • Cybersecurity • Defense
Develops and maintains business intelligence reports, analyzes data, integrates Salesforce and Tableau, and supports catalog and pricing data processes. Partners with sales, business development, and program teams on growth strategies, KPIs, market reports, and strategic planning. Provides reporting-tool expertise, training, and ad hoc market analysis while monitoring commercial aerospace trends. The role is fully remote and requires aerospace industry experience.
Top Skills: AlteryxDeep Learning FrameworksLookerMachine LearningExcelMicrosoft PowerpointMicrosoft WordMs Power QueryPower BIPythonSalesforceSQLTableauTableau PrepTime Series Analysis
27 Minutes Ago
Remote
USA
60-70 Hourly
Entry level
60-70 Hourly
Entry level
Consumer Web • Healthtech • Professional Services • Social Impact • Software
Source and engage technical talent, primarily for engineering roles, using LinkedIn Recruiter, Juicebox, Ashby, referrals, talent communities, market mapping, and creative outreach. Partner with recruiters and hiring managers to refine profiles, share market insights, build diverse pipelines, conduct initial candidate screens, and deliver a positive candidate experience in a fast-paced, metrics-driven environment.
Top Skills: AshbyJuiceboxLinkedin Recruiter
30 Minutes Ago
Remote or Hybrid
169K-281K Annually
Senior level
169K-281K Annually
Senior level
Big Data • Fintech • Information Technology • Business Intelligence • Financial Services • Cybersecurity • Big Data Analytics
Serve as a trusted technical advisor to telecommunications carriers and enterprise customers. Design carrier-grade messaging, voice, identity, and fraud-prevention solutions; translate customer needs into scalable architectures and product enhancements; lead technical workshops and architecture reviews; support RFI/RFP responses; and collaborate with Product, Engineering, Sales, Business Development, and Customer Success teams to drive adoption and successful implementations.
Top Skills: 3GppAtisAWSCarrier Network ArchitectureCi/CdCloud-Native ArchitecturesContainerized ApplicationsDistributed SystemsDockerGsmaIetfInfrastructure AutomationJSONKubernetesMessaging EcosystemsMicroservicesOn-Premises EnvironmentsPrivate CloudPublic CloudRest ApisSignaling FrameworksSipStir/ShakenTelecommunications ProtocolsVoice Communications

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account