Software Engineer, Benchmarking

Remote (United States)

Job Details

Location: United States

Workplace: Remote

Employment Type: Full-Time

Experience: 2+ years of professional software engineering experience

Core Areas: AI benchmarking, evaluation infrastructure, benchmark development, LLM evaluations, Inspect, AI provider integrations, experimentation, research collaboration

Schedule: Flexible hours; overlap with UTC-8 (Pacific Time) and UTC (Greenwich Mean Time) is preferred

Travel: Attendance at three annual team retreats is strongly encouraged

Compensation: $150,000-$325,000 per year

About the Role

This opportunity is for a Software Engineer, Benchmarking focused on building, running, and improving infrastructure used to evaluate frontier AI models. The role combines production software engineering with AI benchmark implementation, evaluation tooling, experimentation, and the development of new benchmarks that help researchers, developers, and policymakers better understand AI capabilities.

You will maintain benchmarking infrastructure, integrate with AI providers, adapt existing benchmarks to internal systems, and help design new evaluations. The position also involves close collaboration with researchers, analysts, and engineers to ensure evaluation data is accurate, useful, and effectively incorporated into research products and publications.

What You'll Do

Implement Benchmarks Develop New Benchmarks Collaborate on Evaluation Research

Qualifications

Required Experience

Required Skills

Preferred Qualifications

Benefits

If you notice a problem with this job, email us at contact@7seventy.net.

Looking for more opportunities?

View All Jobs