
AI research & LLM evaluation
I evaluate, build with and study large language models — bringing a behavioral scientist’s measurement rigor to AI.
Selected work
Model-output evaluation at micro1 (2025–present). As a contract statistician and data scientist, I was selected for quality-control responsibilities after strong AI-refinement performance. I apply calibrated statistical judgment to reviewing frontier-model outputs, detecting errors and evaluating quality. Independently certified by micro1 in 47 skills, including rubric design and quality evaluation, prompt authoring and AI coding agents.
Spark (2026–present). Founder and AI programmer of a fully automated, LLM-based fitness-coaching app — from market discovery to a working web prototype in its first quarter. I engineered its data and content pipeline (exercise catalog, workout protocols, Supabase edge functions and an exercise-figure rendering pipeline with automated QA gates), orchestrating and reviewing AI coding agents (Claude, Cursor, Codex).
Research on learning and control. Computational models of reinforcement learning, motivation and control (eLife, 2025; Neuroscience & Biobehavioral Reviews, 2022), and a manuscript in preparation on patterns of long-term task engagement in large language models.
Toolkit: LLM workflows and evaluation · rubric design · prompt authoring · AI coding agents · Python, R, SQL · statistical modeling and calibration.