PPE undergraduate at Wuhan University
Exchanging at University of Chicago
Applied AI systems 路 LLM evaluation 路 Computational social science 路 Social Psychology
I study what happens between a model's fluent output and a decision that must be justified. My work combines empirical measurement with applied AI systems, especially in policy research and survey-based social science.
A useful system should preserve evidence, expose uncertainty, and remain answerable to human judgment.
| Project | What it asks or builds | Stack |
|---|---|---|
| CGSS Gender-Attitude LLM Audit | Can locally served LLM synthetic respondents preserve human distributions, subgroup heterogeneity, and joint response structure? | Python 路 R 路 local LLMs |
| Environmental Policy Monitoring Agent | A sanitized reference implementation derived from production work on acquisition, normalization, deduplication, traceable analysis, and human review. | Node.js 路 web acquisition 路 LLM workflows |
| reg2paper | A lightweight toolkit for publication-oriented regression tables and model diagnostics. | R 路 modelsummary 路 flextable |
I am developing the CGSS audit into a reusable, estimand-aware evaluation framework for LLM-generated survey responses. The public system separates a runnable synthetic fixture from restricted CGSS microdata and treats model outputs as objects of validation rather than substitutes for respondents.
- Match empirical claims to identification strength.
- Evaluate distributions, heterogeneity, dependence, and failure modes鈥攏ot only means.
- Keep restricted data, production credentials, and stakeholder information outside public repositories.
- Make the reproduction path and its boundary visible.
Other empirical research
Python 路 R 路 tidyverse 路 pandas 路 ggplot2 路 LaTeX 路 Git
