SWE-Bench Task Auditor
Mercor
$70 – $90/hour · as listed
As a SWE-Bench Task Auditor at Mercor, you'll evaluate software engineering benchmark tasks designed to train and test AI models at a frontier lab. Your work involves reviewing repository-level coding challenges, assessing the accuracy of reference solutions, validating test frameworks, and ensuring grading systems function correctly. You'll provide detailed, structured feedback using standardized rubrics to help maintain the integrity of these benchmarking datasets.
This intermediate-level AI data work role suits experienced software engineers with 3+ years of professional coding experience and a track record in open-source projects. You'll work remotely in the United States on an hourly basis, with an expected commitment of around 40 hours per week.
The position offers $70–$90 per hour as listed. This is platform-based work through Mercor's marketplace, ideal for engineers looking to contribute to AI model development while leveraging their deep technical expertise in code quality and assessment.
From the listing: Mercor marketplace · hourly · 40h/week. Evaluate the quality, correctness, and reproducibility of software-engineering benchmark tasks used to train and evaluate a frontier AI lab's models. You'll assess repository-level tasks, reference patches, test harnesses, and grading integrity — and provide clear, rubric-based written feedback. Basic Qualifications • 3+ years professional software engineering • Real open-source contribution or m
Applications are completed on the listing’s own site. Pay is shown as posted by the source (“as listed”) or the aggregator’s estimate where marked — offers and availability change, and individual results vary.
Want help getting selected — and finding the best-paid work for your country? See how membership works →