1 option
Leveraging synthetic data from large language models to steer and enhance model learning Ajay Patel
- Format:
- Book
- Thesis/Dissertation
- Author/Creator:
- Patel, Ajay, author.
- Language:
- English
- Subjects (All):
- Computer science.
- Computer engineering.
- Information science.
- 0984.
- 0464.
- 0723.
- 0800.
- Local Subjects:
- Computer science.
- Computer engineering.
- Information science.
- 0984.
- 0464.
- 0723.
- 0800.
- Genre:
- Academic theses
- Physical Description:
- 1 online resource (235 pages)
- Contained In:
- Dissertations Abstracts International 87-12A
- Place of Publication:
- Ann Arbor : ProQuest Dissertations and Theses, 2026
- Language Note:
- English
- Summary:
- Synthetic data, training data generated by large language models (LLMs), has enabled a promising direction for model improvement. LLMs can generate synthetic data to improve their own performance or enhance smaller models on specific tasks, sometimes even enabling these models to surpass the larger LLM's capabilities. Despite these promises, using synthetic data for model enhancement is counterintuitive and challenging to implement in practice. The characteristics of synthetic data that lead to improved model performance-and the mechanisms by which models learn from it-are poorly understood and vary by task. In this dissertation, I contend that synthetic data's effectiveness hinges on using it as a method of encoding prior knowledge about a task. From this perspective, I reframe synthetic data as an inductive bias-a mechanism through which human practitioners equip models with helpful learning biases and assumptions to improve their performance. LLMs enable practitioners to, at scale, steer and enhance the learning of neural models by allowing them to shape the training examples fed into a model with their prior knowledge about a task. Practitioners can do this by influencing the synthetic data generation, filtering, and selection process. The mechanism for model learning after training on synthetic data derives from this prior knowledge. I first demonstrate this in an in-context learning setting, using a self-consistency prior to bootstrap synthetic examples to improve the machine translation performance of a LLM. Next, I introduce a framework, DataDreamer, to make synthetic data generation and training workflows easier to implement. Using DataDreamer, I demonstrate in a fine-tuning setting how to encode task-specific learning biases in synthetic datasets to enable training stronger embedding models, agent models, and vision-language models. Finally, I use synthetic data to impose an inductive bias during pre-training, the fundamental learning process of LLMs. I restructure pre-training documents into synthetic instruction-answer pairs to perform supervised, instruction-tuning at scale instead of the typical, self-supervised next token prediction. This better aligns the pre-training objective with the distribution of downstream, real-world usage and yields stronger models under equivalent compute budget
- Notes:
- Source: Dissertations Abstracts International, Volume: 87-12, Section: A.
- Advisors: Callison-Burch, Chris Committee members: Roth, Dan; Ungar, Lyle; Watts, Duncan; Hajishirzi, Hannaneh; Raffel, Colin
- Ph.D. University of Pennsylvania 2026
- Vendor supplied data
- Local Notes:
- School code: 0175
- ISBN:
- 9798247974062
- Access Restriction:
- Restricted for use by site license
The Penn Libraries is committed to describing library materials using current, accurate, and responsible language. If you discover outdated or inaccurate language, please fill out this feedback form to report it and suggest alternative language.