I am a Ph.D. candidate at Carnegie Mellon University's Human-Computer Interaction Institute, co-advised by Dr. Ken Koedinger and Dr. Sherry Tongshuang Wu. I work at the intersection of human–AI interaction, generative AI, and the learning sciences.
Large language models are redefining human work, yet people benefit from them unequally because of how they collaborate with AI and how human–AI systems are designed. In my Ph.D., I've been focused on programming, the frontier where LLMs are changing both what it means to program and who gets to do it. My research asks: how could we enable everyone for the AI era? Specifically, I ask:
- What to teach: identifying human skills that stand the test of time as models improve, especially communication and evaluation Not Everyone Wins, CHI'26 pAIr, AIED'23 Workshop
- How to teach: building LLM-based learning systems that train these higher-order thinking skills through deliberate practice HypoCompass, AIED'24 🏆 ROPE, TOCHI'25 ImaginAItion, AIED'26
- How to evaluate: designing systematic, scalable assessments of human learning and human–AI collaboration SPHERE, ACL Findings'25 RECAP, ACL Demo'26 Sim2Real Gap, COLM'26
My work has been published at top venues such as CHI, TOCHI, AIED, ACL, NeurIPS, and COLM, and featured in media such as The New York Times. Tools I built have been used in university classrooms with hundreds of students. The work I led has won the Google Academic Research Award and has been recognized with distinctions including a Best Paper, Honorable Mention, and the Block Center Fellowship on AI & Society. See selected projects below and research for the full publication list.
News
- Jun 2027UpcomingI will serve as Publicity Chair for the Human–AI Interaction Conference (Washington, DC).
- Dec 2026UpcomingOur paper OdysSim will appear at NeurIPS 2026 in Sydney.
- Oct 2026UpcomingOur paper Mind the Sim2Real Gap will appear at COLM 2026 in San Francisco.
- Sep 2026I was named a 2026 Fellow on AI & Society by CMU’s Block Center for Technology and Society.
- Aug 2026Finished my Applied Science internship at Microsoft in Redmond, working with Shamsi Iqbal, Siddarth Suri, and Pranav Khape on how expertise and trust shape AI-assisted work.
- Jul 2026Presented my work “GenAI Defaults to Bias!” and my doctoral consortium project on what calculus we should teach in the age of AI at the Festival of Learning (AIED 2026) in Seoul. I also co-organized the Workshop on AI Literacy Education for All.
Older news
- Jul 2026Two papers accepted to ACL 2026: RECAP (Demos) and What Prompts Don’t Say (Findings).
- Jul 2026Gave an invited talk on What to Teach and How to Teach Programming in the AI Era at UW.
- Jun 2026Gave an invited talk at HKUST (Guangzhou).
- May 2026Gave an invited talk on What to Teach and How to Teach Programming in the AI Era at UIUC.
- Apr 2026Presented my work Not Everyone Wins with LLMs, ROPE (TOCHI), and my doctoral consortium project Training for the Future at CHI 2026 in Barcelona! I also spoke on a panel at the Human-Agent Collaboration workshop and served as a student volunteer.
- Apr 2026Gave invited talks at UChicago, UCI, UCSD, UC Berkeley, and Stanford.
- Mar 2026Gave invited talks at Columbia and University of Michigan.
- Feb 2026Gave invited talks on What to Teach and How to Teach Programming in the AI Era at University of Pennsylvania and Cornell.
- Jan 2026Our proposal What Calculus Should We Teach in the Age of AI? received a $150K LearnVia seed grant. I’m a Co-PI and led the proposal writing.
- Dec 2025Gave invited talks on Training Future-Proof Developers at HKU, HKUST, and ECNU.
- Dec 2025Gave a keynote on workforce preparation for AI literacy at Plug and Play Tech Center, and an invited talk at Pitt on training non-experts to use LLMs for data science.
- Nov 2025Our work on HypoCompass and ROPE was featured in The New York Times.
- Oct 2025Joined the MIT Media Lab’s Benchmarks for Human Flourishing with AI workshop in Boston as an invited expert.
- Oct 2025Gave an invited talk at the Mila HCAI seminar on Training Requirement-Driven LLM Use for Prompt Programming.
- Sep 2025My advisors received a $75K Google gift fund for evaluating LLMs’ pedagogical implementations in real AI tutoring. I’m the lead PhD researcher on the project.
- Aug 2025Gave an invited best-paper presentation of HypoCompass at the IJCAI 2025 Sister Conferences track in Montreal.
- Jul 2025Presented SPHERE, our evaluation card for human–AI systems, at ACL 2025 in Vienna.
- Jul 2025Gave an invited talk at Bloomberg on training non-experts to use LLMs for data science.
- Jun 2025Gave an invited best-paper presentation of HypoCompass for the IAALDE network at CSCL/ISLS 2025 in Helsinki.
- Apr 2025We presented ImaginAItion, our GenAI literacy game, at CHI 2025 in Yokohama.
- Apr 2025ROPE was accepted to ACM TOCHI!
- Nov 2024My advisors received a Google Academic Research Award ($100K) for explicitly training humans toward human–AI collaboration. I helped write the proposal and am the lead PhD researcher on the project.
- Aug 2024Gave an invited talk at Google DeepMind on Scaling Student Support: Online Question-Answering with LLM Tutors.
- Jul 2024🏆 Traveled to Recife, Brazil for AIED 2024, where HypoCompass won the Best Paper and Best Interactive Event awards!
Selected Projects
What to teach
A semester-long classroom study (36 students, 7,000+ logs) showing LLMs don’t close the experience gap in AI-assisted data science: many students miss chances to use AI for explanation and evaluation.
A review contrasting human–human pair programming with human–AI “pAIr” programming, surfacing which collaborative skills should transfer to working with AI.
How to teach
A teachable agent where novices debug AI “learners’” code. It improved debugging performance by 12% and efficiency by 14%, generating practice materials 4× faster than instructors.
A multiplayer prompting game that reveals how GenAI defaults to bias under underspecified prompts. It improved players’ understanding of GenAI behaviors by 35%.
How to evaluate
An evaluation card documenting what, how, who, when, and how-validated for human–AI system evaluations, applied to 39 systems across HCI and NLP.
Running τ-bench with 451 real users and 31 LLM simulators shows simulated users are overly cooperative and uniform, and that stronger general models don’t necessarily simulate users more faithfully.
* equal contribution † mentoringSee all publications →
Beyond Research
Before my Ph.D., I earned dual B.S. degrees in Computer Science and Cognitive Science, with a minor in Design for Learning, from CMU in 2022 (Phi Beta Kappa). Outside research, I like musicals, give tours as an art museum docent, and coach snowboarding. Learn more about my activities, portfolio, and some fun facts!