Siri T.
Group Product Manager, Agentic AI & Servicing @ PayPal | ex-Google (Gemini), Atlassian
A few weeks ago I ran an AI Evals workshop for 70+ PMs at Atlassian. The hands-on parts went fine - building test sets, comparing model variants, reading confusion matrices. The question that kept coming back afterward, in the room and in DMs for weeks, was the one no one had a clean answer to: What do we measure when accuracy isn't the answer? It's the same question that surfaces every time I review a new AI feature. A PM walks me through it. Slide deck, model card, accuracy number on slide three. The rest of the deck is launch logistics. For a while I asked the questions everyone asks — eval set, test methodology, comparison to last quarter's model. The conversation would close. The feature would ship. Sometimes a few months later we'd be in a different room talking about why something broke. I don't ask those questions first anymore. I ask this one: What does the user lose when this is wrong? Wrote up how I think about it — and how Customer Trust at Atlassian shaped the framing — in a new article 👇