AI Evaluators: Assessing A Shopping Assistant
JobgetherΒ·1 day ago
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a AI Evaluators: Assessing A Shopping Assistant based in United States.
This role offers an opportunity to help evaluate and improve the quality of an AI-powered digital shopping assistant.
You will analyze real-world e-commerce interactions to determine whether responses are accurate, logical, useful, and aligned with user needs.
By identifying subtle failures and weaknesses, you will provide structured feedback that directly contributes to improving model performance.
The work combines AI evaluation, quality assurance, e-commerce analysis, and structured data assessment in a practical, research-oriented environment.
You will also develop rubrics and verifiers that create consistent standards for evaluating future AI responses.
This is a sustained remote engagement suited to detail-oriented professionals who enjoy analyzing complex text interactions and shaping better AI experiences.
Accountabilities:
- Review real user interaction traces with an AI-powered shopping assistant, carefully assessing conversations and responses within a dedicated evaluation platform.
- Identify logical failures, factual inaccuracies, irrelevant or unhelpful responses, and poor product recommendations, including subtle issues that may negatively affect the shopping experience.
- Analyze the quality of AI-generated responses from an e-commerce perspective, considering whether recommendations appropriately address user queries and real-world shopping needs.
- Create structured evaluation rubrics that establish clear, repeatable criteria for judging response accuracy, helpfulness, reasoning quality, and overall usefulness.
- Develop verifiers and other structured evaluation mechanisms that can consistently assess future responses and help surface recurring model weaknesses.
- Contribute insights from individual evaluations to broader efforts to improve AI model behavior, response quality, and performance on real-world e-commerce scenarios.
- Maintain a sustained evaluation workload of at least 20 hours per week while working independently and maintaining a high level of accuracy and consistency.
- Experience in data evaluation, quality assurance, AI training, data annotation, software testing, prompt engineering, or a closely related analytical discipline.
- Strong analytical and critical-thinking skills, with the ability to identify subtle logical errors, inaccuracies, inconsistencies, and quality issues within written AI interactions.
- Familiarity with e-commerce search, online shopping journeys, product discovery, recommendations, and digital shopping experiences.
- Ability to analyze complex text interactions in depth and distinguish between technically correct responses and responses that are genuinely useful to the user.
- Experience creating structured evaluation criteria, annotation frameworks, testing methodologies, rubrics, or similar quality-assurance systems is valuable.
- Strong attention to detail and consistency, with the ability to apply evaluation standards objectively across a high volume of interactions.
- Ability to work independently in a remote environment, learn new evaluation tools and processes, and communicate findings clearly.
- Availability to commit to a sustained workload of 20+ hours per week.
- Compensation of $50 USD per hour.
- Fully remote work, providing flexibility to complete evaluation activities from within the United States.
- Sustained part-time engagement requiring 20+ hours per week, allowing for a consistent workload.
- Opportunity to contribute directly to the evaluation and improvement of AI-powered shopping technology.
- Hands-on exposure to AI evaluation, model quality assessment, e-commerce interactions, structured rubrics, and verification frameworks.
- Opportunity to apply expertise in quality assurance, data evaluation, e-commerce, software testing, or AI training to real-world AI development.
Requirements:
Benefits:
Market context
Measured from remote postings we have tracked ourselves β not self-reported survey data.
What Entry Level Data Science roles in Americas pay
- 25th
- $50k
- Median
- $66k
- 75th
- $94k
Based on 537 comparable postings with disclosed salaries, last 12 months.
How Jobgether is hiring
- Last 90 days
- 7,946 roles
- Total tracked
- 25,311
- Hiring across
- 13 job families
Tracked since January 2026.