Member of Technical Staff
Contributed to Grok 4.20, 4.3, 4.5, 4.6, and 4.7. Also worked on RL, SFT, Multi-agents, Cybersecurity training, and Grokipedia.

I studied EECS and Applied Math at UC Berkeley, where I researched speech-language models in the Berkeley Artificial Intelligence Research (BAIR) Lab with Cheol Jun Cho, Nicholas Lee, and Professor Gopala Anumanchipalli, and worked on LLM safety with David Wagner’s Research Group alongside Zhanhao Hu and Professor David Wagner.
When I’m not working, I enjoy soccer and spending time with my four cats.
Contributed to Grok 4.20, 4.3, 4.5, 4.6, and 4.7. Also worked on RL, SFT, Multi-agents, Cybersecurity training, and Grokipedia.
Built Stripe’s first evaluation setup for RAG and tool-calling agents, so teams could see where an agent was failing and fix it. That helped get 8+ agents into production, with about $3M in projected recovery.
Trained models that guess video encoding partitions at 91.9% bitwise accuracy, then used them to cut an SVT-AV1 search down to 10–20% of brute force, with about 5–6% quality loss. The pipelines ran over ~60GB and 10 million examples.
Wrote a Flink job that found degraded Hadoop nodes and turned them off, which took that from 45 minutes down to zero. Pages for bad nodes dropped from 30 a month to 9, and the heuristic caught about 90% of them.
eecs 182 - Designing, Visualizing and Understanding Deep Neural Networks - graded the final and the project, wrote homework and exam questions, and taught two discussions a week.
eecs 127 - Optimization Models in Engineering - graded homework and exams.
eecs 16a - Foundations of Signals, Dynamical Systems, and Information Processing - graded, debugged the homework each week, wrote exam questions, and held office hours.
also on Google Scholar. a star means equal contribution.