Build Your Evals Using AI Experiments on a Real Product Use Case
If you run experiments for a living, you already have the instincts evals demand: hypothesis discipline, metric validation, and healthy skepticism about your own results. This workshop points those instincts at AI outputs.
What you'll walk away with:
A working evals suite built on a real AI use case, taken from first prompt to measured improvement using the same test-and-learn loop you already run.
A failure taxonomy for your AI outputs: the specific ways they break, like hallucination, missing info, and wrong format. Think of it as a research-backed problem inventory, just for a model instead of a page.
A repeatable evaluation loop you can run on any AI feature. If you can design and read an A/B test, you can do this. No AI engineering or data science background needed.
A clear line from the quality metrics evals produce to the outcomes you already own: conversion, retention, adoption.
Rewatch the Keynote
Want to join the next just product conference to see keynotes like this live onsite or in the livestream?
Make sure to get your ticket today
Catalina Turlea
Catalina Turlea is Co-Founder of Lovelaice, an AI evals startup that helps product teams find out whether their AI features work before they ship them.
Before Lovelaice she co-founded nilo.health and spent six years there as CTO. She has been building software for 14 years across mobile, backend and AI, and teaches AI evaluation to product managers on Maven.
Get your ticket for 2026
Save up to 100€ with our late bird ticket

