Loading...
Outlier
Bring your expertise to the problems today’s most advanced AI agents still struggle to solve.
We’re looking for experienced professionals to develop challenging tasks for an AI benchmark. You’ll draw on problems from your own field and turn them into clearly defined challenges that an AI agent must solve from scratch in a Linux terminal.
What you’ll do
Identify difficult, meaningful problems in your area of expertise.Translate those problems into tasks that can be solved in a Linux terminal.Design challenges that expose limitations in advanced AI agents. Each task is tested against three leading models, with three attempts per model, and must prove difficult enough that the models mostly fail. Who we’re looking for
A completed master’s degree or higher. Candidates with a bachelor’s degree, at least 10 years of relevant professional experience and a publication may also be considered.At least 10 years of professional experience in a technical domain.Expertise in coding, machine learning, systems, cybersecurity or hardware.At least one academic or professional publication.Practical comfort working in a Linux terminal.Currently based in the United States. Project details
Task-based project expected to run for at least seven weeks.Onboarding includes two courses and a graded assessment, taking approximately 90 minutes.A live onboarding webinar is also required. Compensation
Earn up to $1,350 per task, with opportunities for higher rates based on task quality and submission volume.Tasks take six to seven hours on average, although initial tasks may take longer as you get familiar with the process.
Apply directly on Braintrust to get started.
Applications beat interviews when you prepare: How to pass AI platform assessments →