My Account Log in

1 option

How (un)Stable Are LLM Occupational Exposure Scores? Evidence from Multi-Model Replication / Michelle Yin, Hoa Vu, Claudia Persico.

NBER Working papers Available online

View online
Format:
Book
Author/Creator:
Yin, Michelle.
Contributor:
Vu, Hoa.
Persico, Claudia.
National Bureau of Economic Research.
Series:
Working Paper Series (National Bureau of Economic Research) no. w35110.
NBER working paper series no. w35110
Language:
English
Physical Description:
1 online resource: illustrations (black and white);
Place of Publication:
Cambridge, Mass. National Bureau of Economic Research 2026.
Summary:
A rapidly growing literature estimates AI's labor-market effects using large language models (LLMs) to self-assess occupational exposure. We demonstrate these measures are highly fragile. Replicating the dominant rubric with three frontier models on identical tasks, we find a 3.6-fold divergence in mean exposure with agreement as low as 57%. This measurement instability alters downstream empirical conclusions: in a difference-in-differences framework, individual-level coefficient magnitudes vary 2.4-fold across annotators, and county level estimates flip from a significant negative to an insignificant positive depending on annotators. We formalize this non-classical measurement error, highlighting the risks of treating evolving LLMs as static instruments.
Notes:
April 2026.
Print version record

The Penn Libraries is committed to describing library materials using current, accurate, and responsible language. If you discover outdated or inaccurate language, please fill out this feedback form to report it and suggest alternative language.

Find

Home Release notes

My Account

Shelf Request an item Bookmarks Fines and fees Settings

Guides

Using the Find catalog Using Articles+ Using your account