← All projects

2022 · Data annotation lead

NLP Data Annotation Pipeline

An annotation pipeline for sentiment and intent labelling that combined automated pre-labelling with structured human review and quality sampling.

Raw text messages flowing through four amber filter panels and emerging as clean labelled records, depicting an NLP annotation pipeline
Raw text moves through ingestion, pre-labelling, human review and validated export before reaching training.
The problem

Manual labelling was slow and inconsistent, and there was no way to measure label quality before data reached training.

Approach
Technologies
Pythonscikit-learnSQLAWSLinux
Outcomes
More work