1 option
Reinforcement Learning from Human Feedback : LLM Alignment and Post-Training.
- Format:
- Book
- Author/Creator:
- Lambert, Nathan.
- Language:
- English
- Subjects (All):
- Reinforcement learning.
- Machine learning.
- Human-computer interaction.
- Physical Description:
- 1 online resource (287 pages)
- Edition:
- 1st ed.
- Place of Publication:
- New York : Manning Publications Co. LLC, 2026.
- Summary:
- This is the authoritative guide for Reinforcement learning from human feedback, alignment, and post-training LLMs. In this book, author Nathan Lambert blends diverse perspectives from fields like philosophy and economics with the core mathematics and computer science of RLHF to provide a practical guide you can use to apply RLHF to your models.
- Contents:
- Reinforcement Learning from Human Feedback
- copyright
- contents
- foreword
- preface
- acknowledgments
- about this book
- about the author
- about the cover illustration
- Part 1. Overview
- 1 Introduction
- 2 A tiny history of RLHF
- 3 Training overview
- Part 2. Core training methods
- 4 Instruction fine-tuning
- 5 Reward modeling
- 6 Reinforcement learning
- 7 Reasoning and inference-time scaling
- 8 Direct-alignment algorithms
- 9 Rejection sampling
- Part 3. Data and preferences
- 10 The nature of preferences
- 11 Preference data
- 12 Synthetic data
- Part 4. Applications and advanced topics
- 13 Tool use and function calling
- 14 Over-optimization
- 15 Regularization
- 16 Evaluation
- 17 Crafting model character and products
- A. Definitions
- B. Beyond "just style"
- C. Practical issues
- references.
- Notes:
- Description based upon print version of record.
- Description based on publisher supplied metadata and other sources.
- Other Format:
- Print version: Lambert, Nathan Reinforcement Learning from Human Feedback
- ISBN:
- 9781638358152
- OCLC:
- 1610751099
The Penn Libraries is committed to describing library materials using current, accurate, and responsible language. If you discover outdated or inaccurate language, please fill out this feedback form to report it and suggest alternative language.