Weekday, Inc. Logo

Weekday, Inc.

QA/Test Engineer

Posted 3 Days Ago
Remote
Hiring Remotely in United States
60-90 Hourly
Junior
Remote
Hiring Remotely in United States
60-90 Hourly
Junior
Validate and harden AI evaluation benchmarks by designing test cases, reviewing tasks and reference solutions, debugging Python evaluation scripts, identifying edge cases and exploits, developing repeatable QA processes, and collaborating with researchers to ensure reliable, reproducible benchmark results.
The summary above was generated by AI

This role is for one of our clients

Compensation: $60-$90 per hour

Join a pioneering AI initiative focused on building the next generation of evaluation benchmarks for frontier AI models. We are seeking experienced QA and Test Engineers to ensure every benchmark is reliable, reproducible, and accurately measures real AI capabilities.

In this role, you will review complex, multi-step evaluation tasks, validate their correctness, identify edge cases, and strengthen testing methodologies before benchmarks are deployed. You'll collaborate closely with AI researchers and task authors to improve evaluation quality, eliminate ambiguity, and ensure benchmark integrity.

This is a fully remote, full-time engagement requiring approximately 35 hours per week.


RequirementsKey Responsibilities
  • Design comprehensive test cases that validate evaluation tasks, including complex edge cases and unexpected scenarios.
  • Review benchmark tasks and reference solutions to identify ambiguity, inconsistencies, missing requirements, and grading gaps.
  • Debug task environments and Python-based evaluation scripts to ensure reliable execution and accurate results.
  • Develop repeatable quality assurance processes, validation checklists, and testing frameworks for benchmark creation.
  • Identify potential shortcuts, exploits, or weaknesses that could compromise evaluation accuracy or benchmark integrity.
  • Collaborate with AI researchers, engineers, and task authors to improve task quality, reproducibility, and technical rigor.
Required Qualifications
  • Master's degree, PhD, or equivalent practical experience in a STEM discipline involving software engineering, research, or advanced technical problem solving.
  • Minimum 1 year of professional experience in Quality Assurance, Test Engineering, Software Engineering, Research Engineering, or a related technical field with strong quality ownership.
  • Proven experience designing test cases, validating complex software systems, and debugging end-to-end workflows.
  • Strong proficiency in Python and Git, with the ability to troubleshoot unfamiliar codebases and technical environments.
  • Excellent analytical thinking, problem-solving skills, and exceptional attention to detail.
  • Experience documenting bugs, test strategies, and technical findings with clear written communication.
  • Experience evaluating AI systems, machine learning models, or AI-generated outputs is preferred.
  • Ability to work independently while managing multiple complex tasks with minimal supervision.
  • Ability to commit approximately 35 hours per week on a consistent basis.
Preferred Qualifications
  • Experience with AI evaluation, benchmark development, or quality assurance for machine learning systems.
  • Background in automation testing, validation frameworks, or software quality engineering.
  • Familiarity with large language models, AI agent workflows, or evaluation pipelines.
  • Experience creating repeatable QA processes for research or engineering projects.
Why Join
  • Play a key role in improving the quality and reliability of next-generation AI evaluation benchmarks.
  • Collaborate with leading AI researchers on cutting-edge evaluation methodologies.
  • Help ensure AI systems are tested against rigorous, real-world scenarios.
  • Apply your software testing and quality engineering expertise to advance AI reliability.
  • Enjoy the flexibility of a fully remote engagement while contributing to impactful AI research.
Equal Opportunity

We are committed to creating an inclusive workplace where all qualified applicants receive equal consideration. Reasonable accommodations are available throughout the application and engagement process.

Contract & Engagement Details
  • Independent contractor engagement.
  • Fully remote with flexible working hours.
  • Expected commitment of approximately 35 hours per week.
  • Project duration may be extended, shortened, or concluded based on project requirements and individual performance.
  • Work does not require access to confidential or proprietary information from any current or former employer.
  • Payments are issued weekly based on approved work completed.
  • At this time, we are unable to support H1-B or STEM OPT candidates.

Similar Jobs

18 Days Ago
Remote
United States
Mid level
Mid level
Information Technology
Lead QA testing for an AWS-hosted government web application and Drupal front end. Design and execute functional, load, performance, failover, security, accessibility (WCAG 2.1 AA / Section 508), and CI/CD smoke tests. Validate IaC (CloudFormation/Terraform), authentication integration, content migration, USWDS theming, cross-browser/responsive behavior, and document defects in Jira/Azure DevOps.
Top Skills: Ams AuthenticationAtoAWSCi/CdCloudFormationCloudfrontDrupal 10Drupal 9Ec2FedrampJIRALambdaRdsS3Section 508SeleniumTerraformUswdsWcag 2.1 AaZephyr
18 Days Ago
In-Office or Remote
Mid level
Mid level
Artificial Intelligence • Big Data • Cloud • Analytics • Business Intelligence • Generative AI • Big Data Analytics
Design, build, and maintain reusable automation test frameworks across UI (Selenium/Cypress/Playwright), API (RestAssured/Postman/Karate), and backend validation. Integrate tests into CI/CD, create containerized test environments, run functional and non-functional tests, support mobile automation, and collaborate with cross-functional teams to define acceptance criteria, conduct reviews, and triage defects.
Top Skills: AllureAppiumAWSAzureBehaveC#CircleCICloudwatchCucumberCypressDockerElkEspressoGatlingGCPGitGithub ActionsGitlab CiJavaJavaScriptJenkinsJmeterKafkaKarateKubernetesNoSQLPactPlaywrightPostmanPythonRabbitMQRestassuredSeleniumSpecflowSplunkSQLTestngTypescriptWebdriverioXcuitest
19 Days Ago
In-Office or Remote
Senior level
Senior level
Information Technology
Lead QA and testing for a DoD software modernization program: plan and execute manual and automated tests (functional, regression, integration, security, performance, accessibility), develop automation frameworks, manage defect lifecycle, support RMF and accreditation, validate WCAG and security findings, report metrics, and collaborate with Agile teams and stakeholders to ensure release quality.
Top Skills: Accessibility Testing ToolsApi TestingAzure DevopsCi/CdDatabase TestingDevsecopsGitlab Ci/CdJenkinsJmeterJunitLoadrunnerPerformance Testing ToolsRest AssuredRmfSeleniumStigWcag

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account