Applied AI engineer transitioning into AI safety research, with a focus on empirical evaluation methodology — sycophancy, honesty, and the measurement infrastructure around alignment claims. Five years at Deepgram building production speech and language systems; before that, federal civic-tech work at NIH on inclusive pandemic data infrastructure during the first weeks of COVID-19. The through-line: what we measure, and how, determines what's knowable about systems — whether the system is a public-health crisis or a frontier model. Currently scoping a capstone project on demographic variation in sycophantic responses as an eval-validity problem.