Blog
Notes from the build
Short, first-person pieces on what actually works when you put AI into production. Evidence over hype, no predictions.
Making model output trustworthy enough to show a customer
The gap between a working extraction demo and a system a business will put in front of a customer is almost entirely validation. Here's how I close it.
02 June 2026
Evaluating LLM outputs with ROUGE, BLEU and BERTScore
You can't ship AI responsibly if you can't say whether it got better. A practical look at the metrics I used to evaluate model output across 20k+ samples.
21 May 2026
Shipping bydhome from zero: what 0→1 actually costs
Co-founding a live marketplace taught me that the AI is the fun 20%. The other 80% is what separates a demo from a product people use.
06 May 2026