← All projects

2023 — Present · AI writing specialist

AI Content Evaluation Framework

A reusable evaluation framework used to rank and critique generative model responses, giving engineering teams a consistent signal for training and model selection.

Grid of amber scoring bars across weighted rubric dimensions, representing a framework for ranking AI-generated content quality
The weighted rubric: factuality, instruction adherence, tone, structure and safety scored on calibrated bands.
The problem

Reviewers were grading model output on instinct, so scores varied wildly between people and were useless as a training signal.

Approach
Technologies
PythonSQLPandasLabel toolingGoogle Workspace
Outcomes
More work