Position Summary
In this remote, hourly contractor role, you will evaluate AI-generated Python code and develop cases that test coding accuracy and reasoning quality. Tasks may include:
Evaluating AI-generated Python code for correctness, efficiency, readability, and adherence to real-world engineering standards
Identifying bugs, logical errors, edge cases, and anti-patterns in AI outputs, and explaining corrections clearly in writing
Applying structured evaluation rubrics to score code quality consistently across varying task types
Developing prompts and test cases that probe AI accuracy across Python programming tasks, including multi-step problems grounded in real codebases
Rating and comparing AI-generated solutions based on correctness, internal consistency, and adherence to the prompt
Profile Requirements
Background in Computer Science, Software Engineering, or a closely related field preferred
Proficiency in Python, including data structures, algorithms, standard libraries, performance optimization, and idiomatic coding practices
Ability to read, write, and critically evaluate Python code across varying complexity levels
Full professional English proficiency
Exceptional attention to detail
Reliable and self-directed, with consistent output quality in a remote, asynchronous workflow
Preferred Experience
2+ years of professional software engineering experience with Python
Prior experience with AI data training, annotation, or model evaluation
About CNTXT AI
CNTXT AI builds artificial intelligence products and data solutions with a focus on making AI accurate, safe, and globally relevant for impact. Our work spans data services, custom AI solutions, and proprietary AI products, with deep expertise in Arabic-native and secure, sovereign solutions.