1 option
Preserving correct behaviors during neural network-based control policy repair Pengyuan Lu
- Format:
- Book
- Thesis/Dissertation
- Author/Creator:
- Lu, Pengyuan, author.
- Language:
- English
- Subjects (All):
- Computer science.
- Computer engineering.
- Electrical engineering.
- Information science.
- 0984.
- 0544.
- 0464.
- 0723.
- Local Subjects:
- Computer science.
- Computer engineering.
- Electrical engineering.
- Information science.
- 0984.
- 0544.
- 0464.
- 0723.
- Genre:
- Academic theses
- Physical Description:
- 1 online resource (176 pages)
- Contained In:
- Dissertations Abstracts International 87-12A
- Place of Publication:
- Ann Arbor : ProQuest Dissertations and Theses, 2026
- Language Note:
- English
- Summary:
- The major challenge for repairing NN-based control policies is the complex relationship between the parameters and the performance metrics. Due to this complexity, it is hard or impossible to identify alternative parameters that improve performance from all initial states. Specifically, when we modify the parameters to produce improved trajectories from some initial states, performance from other initial states will be compromised. Unfortunately, the state-of-the-art literature has yet to address this challenge.We, therefore, formulate the Repair with Preservation (RwP) problem, which aims to repair as many failed initial states as possible while safeguarding the correctness of the previously successful initial states. To conquer this problem, we first propose Incremental Simulated Annealing Repair (ISAR), which uses simulated annealing to update the NN parameters. Upon every parameter update, ISAR preserves the successful initial states using a log-barriered energy function. Case studies on Unmanned Underwater Vehicle from DARPA challenge, OpenAI Gym Mountain Car, and F1/10 Race Car have shown the effectiveness of ISAR in protecting correct behaviors while repairing incorrect ones.One drawback of ISAR is the large computational cost. Specifically, it needs to compute the log-barrier function on every parameter update, imposing a large overhead. To reduce this cost, we design a more sophisticated repair technique: ISAR with Interpolation (ISAR-I). Instead of safeguarding trajectories by expensive log barrier functions, ISAR-I allows them to be compromised, and fixes them after the repair. The fix is done by efficient interpolation on the stability-plasticity trade-off, a subroutine that is also our design, named Imprecise Bayesian Continual Learning (IBCL). Case studies on Unmanned Underwater Vehicle, Mountain Car, and F1/10 Race Car show that ISAR-I maintains the same repair and preservation performance as ISAR, while costing only 6.5%, 19.6%, and 4.1% of computational time
- Notes:
- Source: Dissertations Abstracts International, Volume: 87-12, Section: A.
- Advisors: Lee, Insup; Sokolsky, Oleg Committee members: Mangharam, Rahul; Eaton, Eric; Matni, Nikolai; Ruchkin, Ivan
- Ph.D. University of Pennsylvania 2026
- Vendor supplied data
- Local Notes:
- School code: 0175
- ISBN:
- 9798247980711
- Access Restriction:
- Restricted for use by site license
The Penn Libraries is committed to describing library materials using current, accurate, and responsible language. If you discover outdated or inaccurate language, please fill out this feedback form to report it and suggest alternative language.