ML evaluation · 2025

DataBytes — ML Engineer Intern

Built evaluation frameworks for language-model outputs across 20k+ samples, with deployment to the HuggingFace Hub and cloud.

By the numbers

20k+
samples evaluated
ROUGE·BLEU·BERTScore
evaluation frameworks built

What I did

At DataBytes I worked on the unglamorous-but-essential side of ML: knowing whether the model is actually any good.

  • Built evaluation frameworks using ROUGE, BLEU, and BERTScore to measure model output quality objectively.
  • Ran them across 20,000+ samples, turning “seems fine” into numbers you can defend.
  • Handled deployment to the HuggingFace Hub and cloud, so the work didn’t stop at a notebook.

This is where I learned that evaluation isn’t an afterthought — it’s the thing that lets you ship AI responsibly.