This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for an AI Interaction Evaluator based in the United States.
This is a senior-level contract opportunity focused on evaluating how modern AI coding agents interact with experienced software engineers.
Rather than developing production software, you’ll use your engineering expertise to assess the quality, usefulness, and judgment demonstrated by AI-generated responses.
You’ll evaluate tools such as Codex, Claude Code, and Cursor across realistic software-development scenarios.
The role goes beyond checking whether code is technically correct, focusing on reasoning, explanations, developer guidance, and overall interaction quality.
You’ll apply rigorous engineering judgment to distinguish genuinely strong AI assistance from responses that are merely plausible or syntactically correct.
Your feedback will help establish clearer standards for what excellent AI-assisted development should look like.
The engagement is fully remote, flexible, and designed for experienced engineers who want to contribute directly to the evolution of AI coding tools.
Accountabilities:
- Evaluate AI interactions: Review AI-generated coding interactions end to end and determine whether responses are useful, accurate at a high level, and consistent with strong engineering practices.
- Assess engineering judgment: Evaluate whether coding agents demonstrate sound technical reasoning, appropriate decision-making, and practical engineering judgment rather than simply producing working-looking code.
- Review explanations and reasoning: Assess the quality of preambles, explanations, reasoning, and guidance, identifying whether they genuinely help developers understand and solve problems.
- Measure response quality: Distinguish between different levels of AI response quality and identify the characteristics that separate adequate interactions from exceptional ones.
- Provide actionable feedback: Deliver clear, direct, and opinionated assessments covering what worked, what failed, and what felt misleading, ineffective, or inconsistent with experienced engineering practice.
- Evaluate developer experience: Consider whether an interaction would build trust with an experienced developer, provide useful direction, or instead create confusion and unnecessary work.
- Shape evaluation standards: Help define what “great” looks like when developers work with AI coding agents and AI-first development environments.
- Apply independent expertise: Make subjective but rigorous judgments without needing to execute or deeply inspect every line of generated code.
Requirements:
- Senior engineering background: Staff-, Principal-, or similarly experienced software engineer, or equivalent depth of practical engineering experience.
- Programming expertise: Strong professional background in Python and/or TypeScript/JavaScript, with the ability to quickly understand different software-development approaches.
- AI coding experience: Hands-on experience with modern AI coding tools such as OpenAI Codex, Claude Code, Cursor, or comparable AI-assisted development platforms.
- AI-assisted development knowledge: Deep familiarity with contemporary AI-enabled software-engineering workflows and how developers collaborate with coding agents.
- Strong engineering taste: Ability to recognize whether a response reflects the thinking, communication style, trade-offs, and technical standards expected from an excellent engineer.
- Critical evaluation ability: Comfortable assessing code and technical reasoning at a high level without needing to execute or manually review every implementation detail.
- Communication skills: Able to articulate clear, concise, and well-supported opinions about the strengths and weaknesses of AI-generated interactions.
- High quality standards: A consistently high bar for software engineering, developer experience, clarity, and technical decision-making.
- Preferred experience: Exposure to AI-first IDEs, prompt design, model evaluation, or structured AI evaluation workflows is advantageous.
- Additional advantage: Experience mentoring senior engineers or establishing engineering standards and best practices is a plus.
Benefits:
- Compensation: $100–$200 per hour.
- Flexible schedule: Approximately 10–20 hours per week, allowing you to balance the engagement with other commitments.
- Remote work: Fully remote contract opportunity across eligible locations.
- Short-term engagement: Initial engagement through early May, with the possibility of extension.
- Fast start: Opportunity to begin as soon as possible.
- High-impact work: Help shape how AI coding agents are evaluated and improve the quality of AI-assisted software development.
- Senior-level contribution: Apply years of engineering experience to challenging questions around AI reasoning, developer experience, and engineering quality.
- Streamlined selection process: Take-home evaluation exercise followed by one behavioral interview.
How Jobgether works:
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
#LI-CL1