2023 — Present · AI writing specialist
AI Content Evaluation Framework
A reusable evaluation framework used to rank and critique generative model responses, giving engineering teams a consistent signal for training and model selection.

The problem
Reviewers were grading model output on instinct, so scores varied wildly between people and were useless as a training signal.
Approach
- Defined weighted dimensions: factuality, instruction adherence, tone, structure and safety.
- Wrote calibration examples for every score band so new reviewers could align quickly.
- Ran inter-annotator agreement checks and revised ambiguous criteria until agreement stabilised.
- Standardised ranking exports so preference data flowed straight into training pipelines.
Technologies
PythonSQLPandasLabel toolingGoogle Workspace
Outcomes
- Inter-annotator agreement rose from roughly 0.52 to 0.81 Cohen's kappa.
- Reviewer onboarding time cut from two weeks to three days.
- Adopted as the default rubric across multiple annotation workstreams.
More work