My Account Log in

1 option

Methodological innovations for clinical research improving reliability, data sharing, and agentic ai across the evidence lifecycle Yiwen Lu

Dissertations & Theses @ University of Pennsylvania Available online

View online
Format:
Book
Thesis/Dissertation
Author/Creator:
Lu, Yiwen, author.
Contributor:
University of Pennsylvania. Applied Mathematics and Computational Science., degree granting institution.
Language:
English
Subjects (All):
Biostatistics.
Computer science.
Bioinformatics.
0308.
0984.
0800.
0715.
Local Subjects:
Biostatistics.
Computer science.
Bioinformatics.
0308.
0984.
0800.
0715.
Genre:
Academic theses
Physical Description:
1 online resource (86 pages)
Contained In:
Dissertations Abstracts International 87-12B
Place of Publication:
Ann Arbor : ProQuest Dissertations and Theses, 2026
Language Note:
English
Summary:
Real-world data from electronic health records, claims databases, and other clinical sources offer an unprecedented opportunity to generate evidence about how care affects health. Realizing that opportunity requires confronting a set of persistent methodological challenges: observational studies are vulnerable to systematic bias and efficiency loss from incomplete follow-up; proprietary data environments limit the independent scrutiny that regulatory science demands; and rigorous study design remains inaccessible to many investigators who lack statistical expertise. This dissertation presents three methodological contributions, each targeting a distinct point of failure along the evidence lifecycle.The first contribution, projected calibration, addresses the reliability of real-world evidence. Negative control outcomes are measurements known to be causally unaffected by the treatment of interest, yet subject to the same data collection processes and confounding structure as the primary outcome. Existing methods treat bias correction and efficiency recovery as separate problems, and existing uses of negative controls are largely confined to the fully-observed setting. Projected calibration uses negative control outcomes to adapt to the structure supported by the data: when they capture systematic distortions linked to treatment assignment, the method reduces bias; when they share residual variation with the outcome, it recovers efficiency lost to incomplete follow-up; and when they provide no useful information, the method reverts to the complete-case estimator, preserving an asymptotic no-harm guarantee. Simulation studies show bias reductions of up to 70% and efficiency gains of up to 20%. In a target trial emulation comparing GLP-1 receptor agonists with metformin using Penn Medicine data, projected calibration attenuated the apparent short-term glycemic benefit observed under complete-case analysis, with outcome missingness exceeding 79% for glucose and 88% for HbA1c.The second contribution, TETRiS (TEnsor-Train Enabled Regulatory Post-Marketing Surveillance), addresses data sharing and independent regulatory reanalysis. Post-marketing safety surveillance depends on real-world data, yet patient-level records are often held in proprietary networks that restrict external access. TETRiS enables a data partner to evaluate the log-likelihood surface of a self-controlled design over a multidimensional parameter grid, compress it using tensor-train decomposition, and transmit a compact one-shot package to a regulator. The regulator can then reconstruct primary analyses, prespecified subgroup analyses, and reduced submodel analyses with full numerical fidelity, without accessing individual patient records. In a post-marketing evaluation of myocarditis and pericarditis following mRNA COVID-19 vaccination using TriNetX data, TETRiS reproduced incidence rate ratios and confidence intervals consistent with pooled patient-level analyses across all analytic layers, including in rare-event settings.The third contribution, PowerGPT, addresses access to rigorous study design. Sample size calculation and statistical test selection are foundational to clinical research, yet remain difficult for investigators who lack formal statistical training. PowerGPT is an agent-based system that integrates a large language model with statistical engines and an R-Python interface to automate power analysis through natural language interaction. In a stratified randomized trial at two institutions, PowerGPT improved task completion rates by 10.4 percentage points for test selection and 21.5 percentage points for sample size calculation, improved accuracy in sample size estimation from 55.4% to 94.1%, and reduced average completion time from 9.3 to 4.0 minutes per question.Taken together, these contributions span the evidence lifecycle from study planning to execution to regulatory evaluation, and reflect a common theme: that scalable, methodologically principled tools can make real-world evidence more reliable, more transparent, and more broadly accessible
Notes:
Source: Dissertations Abstracts International, Volume: 87-12, Section: B.
Advisors: Chen, Yong Committee members: Marmarelis, Melina Elpi; Eaton, Eric
Ph.D. University of Pennsylvania 2026
Vendor supplied data
Local Notes:
School code: 0175
ISBN:
9798247979890
Access Restriction:
Restricted for use by site license

The Penn Libraries is committed to describing library materials using current, accurate, and responsible language. If you discover outdated or inaccurate language, please fill out this feedback form to report it and suggest alternative language.

Find

Home Release notes

My Account

Shelf Request an item Bookmarks Fines and fees Settings

Guides

Using the Find catalog Using Articles+ Using your account