|
Pengrui Han (Barry)
I am currently in the MSCS program at UIUC, advised by Prof. Jiaxuan You.
I am also a researcher in the MIT Brain and Cognitive Sciences department, working with Prof.
Evelina Fedorenko in the EvLab.
I received my B.A. in Mathematics and Computer Science from Carleton College, a leading liberal arts college in the US.
During my undergrad, I was fortunate to work with Prof.
Anima Anandkumar in the Anima AI+Science Lab at Caltech.
韩芃睿  / 
Email  / 
Google Scholar  / 
GitHub  / 
LinkedIn  / 
Twitter
|
|
|
Research
My research is broadly driven by questions about intelligence and cognition. As AI systems develop increasingly complex cognitive abilities and behaviors, I am particularly interested in two directions:
(1) Using AI to study intelligence. AI systems provide unique experimental objects for studying the principles underlying intelligence: they exhibit many of the cognitive abilities we have traditionally studied in humans, while their behavior, self-reports, and internal computations can all be observed and experimentally intervened upon.
(2) AI evaluation and safety. As these systems become increasingly capable, I want to develop rigorous ways to evaluate, understand, and monitor their behavior and internal processes, ultimately helping make AI systems safer and better aligned.
My research across these directions often combines behavioral evaluation, mechanistic interpretability, AI alignment, and insights from cognitive science and neuroscience:
|
Characterizing how models reason, generalize, and (mis)align, including studies of alignment, limitations, and trustworthy reasoning.
|
Interpreting model internals to understand the circuits, representations, and algorithms that give rise to intelligent behavior and drive observable performance and failures.
|
Current AI still falls short of biological intelligence in many ways; I aim to build systems that are more adaptive, continuously evolving, and safe.
|
Using AI as a scientific instrument to understand the human mind and tackle open scientific questions in cognition, language, and memory disorders.
|
If any of this resonates with your interests, feel free to reach out and let's connect / collaborate!
|
Selected Publications
|
|
Modular Cognitive Architecture Emerges in Large Language Models
Pengrui Han, Jacob Andreas, Evelina Fedorenko†, and Andrea Gregor de Varda† († Co-senior Authors)
Preprint, 2026
project
/
code
/
manuscript
/
thread
Through circuit analyses across 46 tasks spanning language, formal reasoning, social reasoning, and physical reasoning, we find that LLMs develop a modular cognitive architecture mirroring the human brain: tasks drawing on the same network in humans recruit overlapping neurons in LLMs, whereas tasks drawing on different networks recruit distinct neurons. Modularity may be a fundamental principle of intelligent systems rather than a biological accident.
|
|
|
The Personality Illusion: Revealing Dissociation Between Self-Reports & Behavior in LLMs
Pengrui Han*, Rafal D. Kocielnik*, Peiyang Song, Ramit Debnath, Dean Mobbs, Anima Anandkumar, and R. Michael Alvarez (* Equal Contribution)
International Conference on Machine Learning (ICML), 2026
NeurIPS LAW Workshop, 2025, Best Paper Honorable Mention
arXiv
/
project
/
code
/
media
What LLMs say about themselves does not predict what they do: coherent self-reports come apart from behavioral dispositions. This cautions against taking fluent self-description as evidence of a coherent underlying self-model, and calls for deeper behavioral and mechanistic evaluation in AI alignment and interpretability.
|
|
|
Large Language Model Reasoning Failures
Peiyang Song*, Pengrui Han*, and Noah Goodman (* Equal Contribution)
Transactions on Machine Learning Research (TMLR), 2026, Survey Certificate
arXiv
/
code
/
proceeding
/
media
We present the first comprehensive survey dedicated to reasoning failures in LLMs. By unifying fragmented research efforts, our survey provides a structured perspective on systemic weaknesses in LLM reasoning, offering valuable insights and guiding future research towards building stronger, more reliable, and robust reasoning capabilities.
|
|
|
In-Context Learning May Not Elicit Trustworthy Reasoning: A-Not-B Errors in Pretrained Language Models
Pengrui Han*, Peiyang Song*, Haofei Yu, and Jiaxuan You (* Equal Contribution)
Findings of Empirical Methods in Natural Language Processing (EMNLP), 2024
arXiv
/
code
/
proceeding
Motivated by the crucial cognitive phenomenon of A-not-B errors, we present the first systematic evaluation on the surprisingly vulnerable inhibitory control abilities of LLMs. We reveal that this weakness undermines LLMs' trustworthy reasoning capabilities across diverse domains, and introduce various mitigations.
|
|
|
ChatGPT Based Data Augmentation for Improved Parameter-Efficient Debiasing of LLMs
Pengrui Han*, Rafal Kocielnik*, Adhithya Saravanan,Roy Jiang, Or Sharir,and Anima Anandkumar (* Equal Contribution)
Conference On Language Modeling (COLM), 2024
arXiv
/
code
/
proceeding
We propose a light and efficient pipeline that enables both domain and non-domain experts to quickly generate synthetic debiasing data to mitigate specific or general bias in their models with parameter-efficient fine-tuning.
|
Selected Awards
- Siebel Scholar, Class of 2027 (2026)
- ICML CTB Workshop Best Paper Award (2026)
- TMLR Survey Certification (2026)
- NeurIPS LAW Workshop Best Paper Honorable Mention Award (2025)
- Phi Beta Kappa Honor Society (2025)
- Carleton College Chang-Lan Award (2024)
- Caltech SURF Award (2023)
- Carleton College Dean's List (2023)
|
Selected Media
- Using Eight Billion AI Personas For Psychology Research Has Its Ups And Downs, Forbes, 2026
- In Conversation with Pengrui Han: Opening Up the AI Brain Reveals Functional Specializations Similar to the Human Brain, MIT Technology Review China, 2026
- 'Not how you build a digital mind': How reasoning failures are preventing AI models from achieving human-level intelligence, Live Science, 2026
- Scientists Found AI’s Fatal Flaw—The Most Advanced Models Are Failing Basic Logic Tests, Popular Mechanics, 2026
- New Framework Simplifies the Complex Landscape of Agentic AI, VentureBeat, 2025
- This AI Paper Explains Why Most "Agentic AI" Systems Feel Impressive in Demos and then Completely Fall Apart in Real Use, MarkTechPost, 2025
- Researchers Discover "Personality Illusion" to Reveal a Profound Disconnect Between Language and Behavior in LLMs, MIT Technology Review China, 2025
|
Teaching
- CS 440: Artificial Intelligence, Teaching Assistant @ UIUC, Fall 2026
- CS 440: Artificial Intelligence, Teaching Assistant @ UIUC, Summer 2026
- CS 411: Database Systems, Teaching Assistant @ UIUC, Spring 2026
- CS 512: Data Mining Principles, Teaching Assistant @ UIUC, Fall 2025
- MATH 241: Ordinary Differential Equations, Teaching Assistant @ Carleton College, Fall 2024
- MATH 321: Real Analysis, Teaching Assistant @ Carleton College, Spring 2024
- MATH 232: Linear Algebra, Teaching Assistant @ Carleton College, Spring 2023
- MATH 232: Linear Algebra, Teaching Assistant @ Carleton College, Winter 2023
|
Academic Services
- Reviewer for journals: TMLR.
- Reviewer for conferences: ICLR, ICML, NeurIPS, ACL, COLM, COLING.
- Reviewer for workshops: Re-Align, LLM-Cognition, BehaviorML, LTEDI, INTERPLAY, AI4Math, LatinX, Assessing World Models
|
|
Misc
Outside research, I play classical flute, hit the slopes any chance I can in winter, swim and play badminton year-round, and try to travel somewhere new every few months, usually somewhere with mountains or water.
|
|