Data Scientist - AI Evaluation and Scoring
Job Description and Requirements
Data Scientist - AI Evaluation and ScoringJob Snapshot
Role: Data Scientist - AI Evaluation and Scoring
Location: Dubai, United Arab Emirates
Industry: IT and Services
Function: Science / R&D
Experience: Minimum 5 years
Job Type: Contractor
Position Overview
The Data Scientist - AI Evaluation and Scoring opportunity in Dubai, United Arab Emirates is a remote contractor role within IT and Services offered through YO IT Consulting. The project focuses on assessing whether artificial intelligence systems can complete realistic data science assignments accurately, logically, and to professional industry standards.
Job Details
Country: United Arab Emirates
City: Dubai
Industry: IT and Services
Function: Science / R&D
Salary: 22000-38000
Estimated salary range based on similar jobs in Dubai; please confirm the final offer with the employer.
Gender: Any
Candidate Nationality: Any
Job Type: Contractor
Role Context
Instead of producing conventional analytics deliverables, the Data Scientist will define the standards used to evaluate them.
Clear judgement is essential because every score must be supported by evidence, applied consistently, and capable of being reproduced by another qualified reviewer.
Key Responsibilities
* Examine realistic data science assignments and identify the characteristics of an excellent response.
* Develop task-specific evaluation criteria for analyses, predictive models, dashboards, and experiment readouts.
* Define scoring levels that distinguish complete, partially correct, unsupported, and incorrect work.
* Evaluate AI-generated and human-produced submissions against approved grading frameworks.
* Assign defensible scores and provide a detailed written justification for every assessment.
* Verify analytical logic, statistical validity, metric selection, calculations, and interpretation.
* Assess whether conclusions are supported by the available evidence and communicated appropriately.
* Review SQL and Python-based analysis for methodological quality and technical correctness.
* Evaluate experiment designs, control groups, success measures, and A-B testing conclusions.
* Check dashboards and reporting outputs for accuracy, usability, and business relevance.
* Apply evaluation standards consistently across repeated tasks and different submission types.
* Identify ambiguous criteria and recommend refinements that improve scoring reliability.
* Incorporate structured comments from senior reviewers and revise work within agreed timelines.
* Participate in calibration exercises to align judgement with other subject-matter experts.
* Maintain concise records explaining decisions, assumptions, and changes made during review.
* Protect project information and follow all contractor confidentiality requirements.
Ideal Profile
Candidates should have at least five years of professional data science experience gained in an industry environment. A background in product, growth, or business-operations analytics within a high-performing technology organization is particularly relevant.
The role requires advanced knowledge of experiment design, A-B testing, metric development, SQL, Python, and the communication of technical findings to executive stakeholders. Exceptional written English, careful attention to evidence, and comfort with peer calibration are essential.
Previous involvement in AI evaluation, model training, data annotation, human-feedback programmes, or similar projects would be advantageous.
Skills Set
* Data science evaluation
* AI output assessment
* Rubric and grading design
* Evidence-based scoring
* Statistical analysis
* Experiment design
* A-B testing
* Metric definition
* SQL
* Python
* Predictive model review
* Dashboard evaluation
* Business operations analytics
* Product analytics
* Growth analytics
* Executive recommendations
* Technical writing
* Quality calibration
* Peer-review feedback
* Analytical reasoning
Why Join Us
This project offers experienced data scientists a different way to apply their expertise by shaping how advanced AI systems are measured and improved. The remote structure provides schedule flexibility, weekly payment through Stripe or Wise, and exposure to emerging AI-evaluation methods without requiring a conventional product-delivery workload.
About the Company
YO IT Consulting is supporting recruitment for a project led by an AI research organization. The engagement brings experienced data science professionals into structured evaluation work designed to improve the quality and reliability of real-world AI capabilities.



