ML evaluation · 2025
DataBytes — ML Engineer Intern
Built evaluation frameworks for language-model outputs across 20k+ samples, with deployment to the HuggingFace Hub and cloud.
By the numbers
20k+
samples evaluated
ROUGE·BLEU·BERTScore
evaluation frameworks built
What I did
At DataBytes I worked on the unglamorous-but-essential side of ML: knowing whether the model is actually any good.
- Built evaluation frameworks using ROUGE, BLEU, and BERTScore to measure model output quality objectively.
- Ran them across 20,000+ samples, turning “seems fine” into numbers you can defend.
- Handled deployment to the HuggingFace Hub and cloud, so the work didn’t stop at a notebook.
This is where I learned that evaluation isn’t an afterthought — it’s the thing that lets you ship AI responsibly.