Authored by: Dr. Joseph Geraci & Dr. Luca Pani
Abstract:
FDA public communications indicate a movement toward agency-wide AI tooling and consolidated data infrastructure, including Elsa 4.0 integrated with HALO, which FDA described as consolidating more than 40 application/submission systems into an AI-accessible environment. The practical implication for sponsors is not that a specific undisclosed algorithm will adjudicate submissions, but that future review can more readily compare claims horizontally across submissions, adverse-event histories, trial populations, endpoints, and precedent. This paper proposes a complementary sponsor-side framework: vertical decomposition of heterogeneous clinical-trial populations into stable, interpretable Model-Derived Subgroups (MDS), followed by a Regulatory Defensibility Index (RDI) that quantifies whether subgroup, biomarker, endpoint, or predictive-enrichment claims are likely to withstand AI-native regulatory interrogation. The proposed RDI aggregates five scored dimensions: Subgroup Stability Score, Endpoint Fragility Score, Placebo Structure Overlap, Mechanistic Coherence Score, and Evidence Calibration Score. Each dimension maps to a regulatory or methodological requirement: prospective enrichment logic, context-of-use credibility, estimand alignment, statistical robustness, multiplicity control, and causal/sensitivity analysis. The document reframes three public case studies – donanemab/TRAILBLAZER-ALZ 2, aticaprant, and ulotaront/DIAMOND – as illustrative defensibility profiles rather than as new efficacy analyses.
Table of Contents
1. Executive summary
2. Scope and evidentiary posture
3. Regulatory context: Elsa/HALO and AI credibility
4. Vertical defensibility: MDS as a sponsor-side complement to horizontal review
5. Regulatory Defensibility Index: five-component metric family
6. Proposed NetraAI Regulatory Defensibility Engine architecture
7. Public case-study defensibility profiles
8. Disease Structure Libraries as defensibility-ready knowledge assets
9. Risk register and self-audit
10. Claim-to-reference map
11. References
Executive summary
The strategic premise of this framework is designed to address an evolving regulatory landscape: FDA has publicly described Elsa 4.0 as an agency-wide internal AI tool integrated with HALO, a consolidated data platform linking more than 40 application/submission systems. FDA also stated that subject matter experts remain responsible for verifying inputs, processes, and outputs. Therefore, the defensibility problem for sponsors should be framed as readiness for more systematic review interrogation, not as speculation about autonomous regulatory decision-making. [1,2]
FDA draft guidance on AI for drug and biological product decision-making emphasizes a risk-based credibility assessment tied to a specific question of interest and context of use. FDA enrichment guidance defines enrichment as prospective selection of patients in whom detection of a drug effect is more likely, and specifically recognizes predictive enrichment as selecting patients more likely to respond to treatment. Together, these documents make the evidentiary target clear: an AI-derived subgroup should not be asserted as a discovery artifact alone; it should be tied to a context of use, prespecified analysis logic, error control, transparent data lineage, and independent validation where possible. [2,3]
NetraAI is positioned here as a vertical decomposition engine: it seeks to isolate stable, explainable substructures inside heterogeneous trial populations, usually through compact variable bundles that can be translated into feasible patient-selection or stratification hypotheses. The framework intentionally avoids claiming that every patient can be explained; instead, it distinguishes explainable subgroups from residual heterogeneity, a posture that is more defensible for small, high-dimensional clinical-trial datasets. Peer-reviewed NetraAI and related Geraci/NetraMark publications provide methodological precedent for explainable subpopulation discovery in depression, ALS, bipolar disorder, and digital-health survey contexts. [10,11,12,13]
The Regulatory Defensibility Index (RDI) proposed in this paper is not a regulatory endpoint and should not be presented as confirmatory evidence. It is a sponsor-side stress-test: a quantitative readiness score that asks whether a subgroup, biomarker, endpoint, or enrichment claim has enough stability, mechanistic coherence, statistical calibration, and placebo separation to merit prospective testing or formal FDA engagement. Its value lies in forcing weak claims to fail before they become expensive protocol commitments or fragile submission narratives.
Key Takeaways

Scope and Evidentiary Posture
This paper synthesizes methodological frameworks for regulatory defensibility, anchoring all foundational claims to vetted public sources. These include official FDA and ICH guidance documents, ClinicalTrials.gov registry records, official sponsor topline disclosures, and PubMed/DOI-indexed literature.
To illustrate how the proposed defensibility framework reasons in practice, the included case studies utilize publicly available aggregate statistics. Because patient-level datasets are not employed in this conceptual demonstration, this document does not attempt to recompute de novo subgroup membership, treatment effects, bootstrap stability, holdout replication, or conditional average treatment effect (CATE) estimates.
The evidentiary posture of this manuscript is intentionally conservative. The appropriate next phase of validation for this framework will involve integrating reproducible code for all public-data-derived values, maintaining a strict separation between public evidentiary demonstrations and confidential, internal NetraAI validations.
Evidence Hierarchy Used in this Draft

Regulatory Context: Elsa/HALO and AI Credibility
FDA announced on May 6, 2026 that it had launched Elsa 4.0 for FDA staff and integrated it with HALO, consolidating more than 40 application/submission systems and portals across centers into a unified environment. FDA described the result through the statement that Elsa now sits on top of FDA data. FDA also stated that the system was deployed with security controls and that human subject matter experts verify inputs, processes, and outputs. [1]
The sponsor-relevant conclusion is not that Elsa/HALO has a known public scoring rubric for subgroup claims. The defensible conclusion is that FDA has publicly moved toward AI-supported internal review infrastructure capable of more efficient horizontal retrieval and cross-document comparison. A sponsor initiated paper should therefore focus on claim traceability, cross-study consistency, and robustness under alternate analytic choices rather than on unsupported assertions about proprietary FDA workflows.
The January 2025 FDA draft guidance on AI in drug and biological product regulatory decision-making provides the most direct template for sponsor-side AI credibility. It asks sponsors to define the question of interest, define the AI model context of use, assess model risk, develop and execute a credibility plan, document results and deviations, and determine adequacy for the context of use. [2]
FDA enrichment guidance is equally central. Predictive enrichment depends on identifying patients more likely to respond to the drug and requires attention to the marker-negative population, Type I error allocation, study design, and labeling implications. The RDI framework proposed below should therefore be read as a pre-enrichment stress-test that informs whether a subgroup should proceed to prospective alpha-controlled validation. [3]
Digital health technology guidance adds a parallel lesson for DHT- or remote-data-derived variables: tools and measurements should be fit-for-purpose, verified/validated, interpretable for the clinical investigation, and accompanied by statistical and risk-management considerations. When NetraAI uses DHT-derived endpoints or features, those upstream data-quality requirements become part of the defensibility chain. [4]
AI-native Regulatory Stress Vectors and Sponsor-Side Countermeasures


Vertical Defensibility: MDS as a Sponsor-Side Complement to Horizontal Review
The strategic contrast is framed here as horizontal FDA interrogation versus vertical NetraAI population resolution. This serves as a conceptual model while avoiding overclaiming.
Model-Derived Subgroups (MDS) are useful only if they survive a higher evidentiary standard than ordinary post-hoc subgroup discovery. The proposed standard is: (1) the subgroup is defined by a feasible and interpretable variable bundle; (2) the endpoint and estimand are prespecified or retrospectively reconstructed with transparent limitations; (3) performance replicates in holdout, resampling, or independent cohorts; (4) the treatment-control or drug-placebo separation is not dominated by placebo behavior; (5) the subgroup is mechanistically coherent; and (6) the evidence remains interpretable after multiplicity and sensitivity analyses.
This orientation aligns with the public NetraAI literature, which emphasizes explainability, small sample/high-dimensional settings, and discovery of clinically meaningful subpopulations, but the external validation requirement remains critical. A subgroup derived from a completed trial is a hypothesis-generating asset until it is prospectively evaluated or independently replicated. [10,11,12,13,24,25]
Regulatory Defensibility Index: Five-Component Metric Family
The RDI is proposed as a five-component sponsor-side readiness and stress-test framework. Each
component is independently scored on a 0–100 scale, and the framework’s output is the full five-
component vector RDI = {SSS, EFS, PSO, MCS, ECS} rather than a single composite. This design is
deliberate: a weighted sum can hide a fatal weakness in any one dimension behind strength in the others, which is precisely the failure mode regulatory review is designed to detect. Where summary judgment is required, the framework adopts a weakest-link rule: the MDS advances only if every component clears its preregistered threshold, rather than additive aggregation.
Default reporting structure: the five-component score vector, the binary pass/fail vector against
preregistered thresholds, and, for ECS specifically, a continuous calibration index showing how close
the binding gate sits to failure.
Component definitions


5.1 Subgroup Stability Score (SSS)
Feasibility: NetraAI’s persona generation produces explicit variable bundles and patient membership every run. Therefore, SSS reduces to a frequency-count problem across the B = 1000 bootstrap, subsample, or seed-perturbation runs. By counting how often each variable (and each pairwise or triplet co-occurrence) appears in the bundle defining the MDS, and how often the same patients land in the same persona, you can normalize the results by the total run count. The result provides per-variable, per-bundle, and per-membership recurrence frequencies derived directly from existing platform output.
Recommendation: Report SSS as a tiered score to give sponsors a defensible answer to whether this is the mathematically identical subgroup across perturbations, without collapsing the data into a single number that hides the underlying structure. We recommend reporting the following three tiers:
- Marginal Frequency: The frequency of each variable in the MDS bundle across the B runs.
- Co-occurrence Frequency: The frequency of variable pairs and triplets evaluated using a lift statistic. This distinguishes variables that appear often merely because every bundle contains them from variables that appear together more frequently than chance would predict. The lift equations for pairs and triplets are defined as:
Lift(X,Y) = P(X ∩ Y) / [P(X) × P(Y)]
Lift(X,Y,Z) = P(X ∩ Y ∩ Z) / [P(X) × P(Y) × P(Z)] - Bundle-Level Stability: The fraction of runs whose top-k variables match the reference bundle at a Jaccard index ≥ 0.7.
5.2 Endpoint Fragility Score (EFS)
Feasibility: once NetraAI has extracted the MDS, fragility computation is a downstream statistical operation on the subgroup’s outcome data and requires no new platform capability; the NetraAI-specific question is whether to compute fragility on the PNR/TR persona, the full responder MDS, or both. Recommendation: compute fragility within each NetraAI-derived persona separately rather than on a pooled subgroup, because the platform’s whole value proposition is that personas are differentiable response phenotypes; pooling them throws away exactly the signal NetraAI was used to find.
For binary endpoints, fragility is quantified using the Fragility Index (FI), defined as the minimum number of patient outcomes that must change from a non-event to an event to render a statistically significant result non-significant. An FI of 0 indicates that a finding is perfectly fragile, often relying on permissive alpha thresholds or lacking robust outcome separation.
5.3 Placebo Structure Overlap (PSO)
Feasibility: NetraAI’s existing PNR/TR versus PR/TNR framing is essentially purpose-built for this component: the platform already distinguishes treatment responders from placebo responders at the persona level, which is the exact decomposition PSO requires. Recommendation: define PSO as the ratio of TR-persona effect to PR-persona effect within the candidate MDS, and use NetraAI’s existing explainability layer to identify which baseline variables drive placebo-response membership so the sponsor can see why a subgroup is or is not placebo-confounded, not just that it is.
5.4 Mechanistic Coherence Score (MCS)
Feasibility: NetraAI’s LLM-augmented explainability layer already generates interpretable narratives of the variable bundles defining each persona, which gives the MCS expert panel a clean, blinded artifact to score against a predefined rubric; the platform supplies the input, while humans supply the judgment. Recommendation: deliver NetraAI’s persona descriptions to a 3+ member expert panel without revealing which platform produced them or which arm showed effect, score against a preregistered rubric, and report Krippendorff’s alpha. Low inter-rater agreement is itself the finding and should be reported as such, not adjudicated away. There is the possibility of passing this task to a specialized AI agent trained for this task with a specialized background in disease etiology.
5.5 Evidence Calibration Score (ECS)
Feasibility: this is the integrative scoring layer that consumes the outputs of the gauntlet – bootstrap stability passes, CATE estimates within the MDS, persona-level p-values, effect sizes, and sample-size adequacy – and asks whether the surviving MDS is jointly defensible across all of them, not just any one in isolation. Recommendation: define ECS as a weakest-link composite: the surviving MDS must clear preregistered thresholds on each of {bootstrap stability, CATE positivity with bounded CI, p-value adjusted for multiplicity, effect size above clinical-significance floor, post-hoc power above 80%}. Report both the binary pass/fail vector and a continuous calibration index, such as the geometric mean of normalized distances above each threshold, so sponsors see exactly which gate is binding and how close to failure the MDS sits, which is the information they actually need for the EOP2 conversation.
Proposed NetraAI Regulatory Defensibility Engine Architecture
This paper proposes a product extension from NetraAI subgroup discovery toward a Regulatory Defensibility Platform. It outlines a defined architecture that can be documented, validated, and discussed with regulators as a decision-support workflow rather than as a black-box decision-maker.

Context-of-use Template for an RDI-Supported NetraAI Analysis
Question of interest: Which baseline-measurable patient subgroup, if any, is more likely to show clinically meaningful benefit from the investigational therapy than the unselected population?
Context of use: Exploratory analysis of completed Phase 2/3 data to generate and prioritize predictive-enrichment hypotheses for prospective validation in a subsequent trial. The model output does not establish efficacy or safety and does not substitute for a prespecified confirmatory analysis.
Model risk: Moderate to high when outputs influence Phase 3 population selection, sample size, SAP alpha allocation, or labeling strategy. Risk mitigation requires transparent documentation, sensitivity analyses, independent validation where feasible, and early FDA engagement. [2,3]
7. Public Case-Study Defensibility Profiles
These case studies are illustrative. They show how RDI reasoning would stress-test public evidence, not how NetraAI would score patient-level data. Patient-level RDI scoring requires access to locked datasets, pre-specified scoring code, and audit logs.
7.1 Donanemab / TRAILBLAZER-ALZ 2
TRAILBLAZER-ALZ 2 was a randomized clinical trial of donanemab in early symptomatic Alzheimer disease with amyloid and tau pathology. Publicly reported JAMA results included 1736 participants. In the low/medium tau population, least-squares mean iADRS change at 76 weeks was -6.02 for donanemab versus -9.27 for placebo; in the combined population, the corresponding values were -10.19 versus -13.11. These differences were statistically significant and support the view that the low/medium tau population is an evidence-relevant subgroup. [26,27]
The defensibility tension is safety. JAMA reported amyloid-related imaging abnormality-edema/effusion (ARIA-E) in 24.0% of donanemab-treated participants versus 2.1% of placebo participants. A sponsor-side RDI should therefore avoid separating efficacy defensibility from safety defensibility. A subgroup can have a stronger efficacy profile while still requiring a safety-aware score that penalizes robust, genotype-linked, or clinically serious adverse-event structure. [26]
Defensibility interpretation: public evidence supports a relatively strong efficacy-oriented profile for pre-specified amyloid/tau-defined populations, but the overall claim must be scored with an explicit ARIA risk module, marker-negative/generalizability discussion, and benefit-risk framing.
7.2 Aticaprant adjunctive therapy in MDD
The phase 2 randomized, double-blind, placebo-controlled aticaprant study in MDD reported efficacy signals for adjunctive kappa-opioid receptor antagonism, including analyses in patients with anhedonia. Publicly available article metadata identifies Neuropsychopharmacology 2024;49(9):1437-1447, DOI 10.1038/s41386-024-01862-x, PMID 38649428. [28]
The high-anhedonia signal is biologically plausible but statistically fragile under conventional confirmatory expectations. This is precisely where RDI is useful: MCS may be moderate because anhedonia is mechanistically relevant, while ECS and EFS can remain weak when the analysis depends on permissive one-sided alpha, lack of multiplicity correction, or responder approximations with FI=0.
A later sponsor statement said the VENTURA phase 3 program was discontinued because of insufficient efficacy, while also stating no new safety signals were observed. This does not negate the phase 2 biological hypothesis; it does indicate that a regulatory defensibility framework should aggressively downgrade nominal, subgroup-sensitive, or multiplicity-sensitive signals before they become Phase 3 commitments. [29,30]
7.3 Ulotaront / DIAMOND-1 and DIAMOND-2
Sumitomo Pharma and Otsuka announced in July 2023 that DIAMOND-1 and DIAMOND-2 did not meet their primary endpoint for ulotaront in acutely psychotic adults with schizophrenia. The official topline disclosure remains the primary public source for program status; ClinicalTrials.gov records provide trial identifiers and registry context. [31,32,33]
Ulotaront serves as a structural no-go counterexample: when placebo improvement is large and treatment-placebo separation is small or inconsistent, post-hoc subgroup claims should be penalized heavily. In the proposed RDI framework this manifests through low or zero Endpoint Fragility Score, poor Placebo Structure Overlap, and weak Evidence Calibration unless a subgroup prospectively shows independent separation.
Defensibility interpretation: a failed whole-cohort primary endpoint does not logically preclude every subgroup hypothesis, but it raises the evidentiary burden. Any MDS claim emerging from such a program would need independent replication, prospective alpha-controlled testing, and explicit placebo-structure modeling before being considered actionable.

8. Disease Structure Libraries as Defensibility-Ready Knowledge Assets
A Disease Structure Library (DSL) is proposed as an indication-specific repository of subgroup structures, endpoint behaviors, placebo patterns, safety liabilities, biomarker logic, and negative-control findings. Its purpose is not to declare a disease taxonomy as final. Its purpose is to preserve what has already been learned about where signal survives and where it fails.
For major depressive disorder, a DSL would include responder/nonresponder structures, placebo-response phenotypes, anhedonia-related hypotheses, dissociation or neuroimaging-linked signatures, and evidence from ketamine/TRD and aticaprant programs. For Alzheimer disease, a DSL would include amyloid/tau staging, cognitive/functional endpoints, ARIA risk structures, APOE-related risk context, and lessons from donanemab, lecanemab, and aducanumab. For schizophrenia, a DSL would include PANSS and functional endpoint behavior, antipsychotic response heterogeneity, placebo-response periods, and lessons from CATIE and ulotaront.
The DSL becomes defensibility-ready only when each entry is tagged by evidence type, replication status, mechanism, endpoint relevance, population feasibility, safety interaction, and RDI component profile. Such a library gives sponsors a reusable substrate for trial design and allows future Elsa/HALO-style review questions to be answered with traceable evidence rather than narrative reconstruction.
9. Risk Register and Self-Audit

10. Claim-to-Reference Map


11. References
1. Food and Drug Administration. FDA Expands AI Capabilities and Completes Data Platform Consolidation [Internet]. Silver Spring (MD): FDA; 2026 May 6 [cited 2026 May 13]. Available from: https://www.fda.gov/news-events/press-announcements/fda-expands-ai-capabilities-and-
completes-data-platform-consolidation
2. Food and Drug Administration. Considerations for the Use of Artificial Intelligence to Support Regulatory Decision-Making for Drug and Biological Products: Guidance for Industry and Other Interested Parties. Draft guidance [Internet]. Silver Spring (MD): FDA; 2025 Jan [cited 2026 May 13]. Available from: https://www.fda.gov/regulatory-information/search-fda-guidance-documents/considerations-use-artificial-intelligence-support-regulatory-decision-making-drug-and-biological
3. Food and Drug Administration. Enrichment Strategies for Clinical Trials to Support Determination of Effectiveness of Human Drugs and Biological Products: Guidance for Industry [Internet]. Silver Spring (MD): FDA; 2019 Mar [cited 2026 May 13]. Available from:
https://www.fda.gov/media/121320/download
4. Food and Drug Administration. Digital Health Technologies for Remote Data Acquisition in Clinical Investigations: Guidance for Industry, Investigators, and Other Stakeholders [Internet]. Silver Spring (MD): FDA; 2023 Dec [cited 2026 May 13]. Available from:
https://www.fda.gov/media/155022/download
5. Food and Drug Administration. Artificial Intelligence and Medical Products: How CBER, CDER, CDRH, and OCP are Working Together [Internet]. Silver Spring (MD): FDA; 2024 Mar [cited 2026 May 13]. Available from: https://www.fda.gov/media/177030/download
6. Food and Drug Administration. Innovative Science and Technology Approaches for New Drugs (ISTAND) Program [Internet]. Silver Spring (MD): FDA [cited 2026 May 13]. Available from: https://www.fda.gov/drugs/drug-development-tool-ddt-qualification-programs/innovative-
science-and-technology-approaches-new-drugs-istand-program
7. International Council for Harmonisation of Technical Requirements for Pharmaceuticals for Human Use. ICH E9(R1) addendum on estimands and sensitivity analysis in clinical trials to the guideline on statistical principles for clinical trials [Internet]. Geneva: ICH; 2019 Nov 20 [cited 2026 May 13]. Available from: https://database.ich.org/sites/default/files/E9-R1_Step4_Guideline_2019_1203.pdf
8. International Council for Harmonisation of Technical Requirements for Pharmaceuticals for Human Use. ICH E17: General principles for planning and design of multi-regional clinical trials [Internet]. Geneva: ICH; 2017 Nov 16 [cited 2026 May 13]. Available from:
https://database.ich.org/sites/default/files/E17EWG_Step4_2017_1116.pdf
9. International Council for Harmonisation of Technical Requirements for Pharmaceuticals for Human Use. ICH E14: The clinical evaluation of QT/QTc interval prolongation and proarrhythmic potential for non-antiarrhythmic drugs [Internet]. Geneva: ICH; 2005 May 12 [cited 2026 May 13]. Available from: https://database.ich.org/sites/default/files/E14_Guideline.pdf
10. Geraci J, Qorri B, Tsay M, Cumbaa C, Leonczyk P, Alphs L, Ballard ED, Zarate CA Jr, Pani L. Explainable AI-driven precision clinical trial enrichment: demonstration of the NetraAI platform with a phase II depression trial. npj Digit Med. 2025;8:749. doi:10.1038/s41746-025-
02143-7. PMID:41360997.
11. Geraci J, Bhargava R, Qorri B, Leonchyk P, Cook D, Cook M, Sie F, Pani L. Machine learning hypothesis-generation for patient stratification and target discovery in rare disease: our experience with Open Science in ALS. Front Comput Neurosci. 2024;17:1199736.
doi:10.3389/fncom.2023.1199736. PMID:38260713.
12. Choi J, Bodenstein DF, Geraci J, Andreazza AC. Evaluation of postmortem microarray data in bipolar disorder using traditional data comparison and artificial intelligence reveals novel gene targets. J Psychiatr Res. 2021;142:328-36. doi:10.1016/j.jpsychires.2021.08.011.
PMID:34419753.
13. Gallucci D, Ho ECY, Geraci J, Loren J, Pani L. Respondents of health survey powered by the innovative NURO app exhibit correlations between exercise frequencies and diet habits, and between stress levels and sleep wellness. Front Psychiatry. 2022;13:945780.
doi:10.3389/fpsyt.2022.945780. PMID:36159919.
14. Walsh M, Srinathan SK, McAuley DF, Mrkobrada M, Levine O, Ribic C, Molnar AO, Dattani ND, Burke A, Guyatt G, Thabane L, Walter SD, Pogue J, Devereaux PJ. The statistical significance of randomized controlled trial results is frequently fragile: a case for a Fragility Index. J
Clin Epidemiol. 2014;67(6):622-8. doi:10.1016/j.jclinepi.2013.10.019. PMID:24508144.
15. VanderWeele TJ, Ding P. Sensitivity analysis in observational research: introducing the E-value. Ann Intern Med. 2017;167(4):268-74.
doi:10.7326/M16-2607. PMID:28693043.
16. Simonsohn U, Simmons JP, Nelson LD. Specification curve analysis. Nat Hum Behav. 2020;4:1208-14. doi:10.1038/s41562-020-0912-z.
17. Wager S, Athey S. Estimation and inference of heterogeneous treatment effects using random forests. J Am Stat Assoc. 2018;113(523):1228-42. doi:10.1080/01621459.2017.1319839.
18. Athey S, Imbens G. Recursive partitioning for heterogeneous causal effects. Proc Natl Acad Sci U S A. 2016;113(27):7353-60.
doi:10.1073/pnas.1510489113. PMID:27382149.
19. Breiman L. Random forests. Mach Learn. 2001;45:5-32. doi:10.1023/A:1010933404324.
20. Collins GS, Moons KGM, Dhiman P, Riley RD, Beam AL, Van Calster B, et al. TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ. 2024;385:e078378. doi:10.1136/bmj-2023-078378.
PMID:38626948.
21. Vasey B, Nagendran M, Campbell B, Clifton DA, Collins GS, Denaxas S, Denniston AK, et al. Reporting guideline for the early-stage clinical evaluation of decision support systems driven by artificial intelligence: DECIDE-AI. Nat Med. 2022;28(5):924-33. doi:10.1038/s41591-022-
01772-9. PMID:35585198.
22. Liu X, Cruz Rivera S, Moher D, Calvert MJ, Denniston AK, SPIRIT-AI and CONSORT-AI Working Group. Reporting guidelines for clinical trial reports for interventions involving artificial intelligence: the CONSORT-AI extension. BMJ. 2020;370:m3164. doi:10.1136/bmj.m3164.
23. Cruz Rivera S, Liu X, Chan AW, Denniston AK, Calvert MJ, SPIRIT-AI and CONSORT-AI Working Group. Guidelines for clinical trial protocols for interventions involving artificial intelligence: the SPIRIT-AI extension. BMJ. 2020;370:m3210. doi:10.1136/bmj.m3210.
24. Kent DM, Rothwell PM, Ioannidis JPA, Altman DG, Hayward RA. Assessing and reporting heterogeneity in treatment effects in clinical trials: a proposal. Trials. 2010;11:85. doi:10.1186/1745-6215-11-85. PMID:20704705.
25. Sun X, Briel M, Walter SD, Guyatt GH. Is a subgroup effect believable? Updating criteria to evaluate the credibility of subgroup analyses.
BMJ. 2010;340:c117. doi:10.1136/bmj.c117.
26. Sims JR, Zimmer JA, Evans CD, Lu M, Ardayfio P, Sparks J, et al. Donanemab in early symptomatic Alzheimer disease: the TRAILBLAZER-
ALZ 2 randomized clinical trial. JAMA. 2023;330(6):512-27. doi:10.1001/jama.2023.13239. PMID:37459141.
27. ClinicalTrials.gov. A study of donanemab (LY3002813) in participants with early symptomatic Alzheimer disease (TRAILBLAZER-ALZ 2). Identifier NCT04437511 [Internet]. Bethesda (MD): National Library of Medicine [cited 2026 May 13]. Available from:
https://clinicaltrials.gov/study/NCT04437511
28. Schmidt ME, Kezic I, Popova V, Melkote R, Van Der Ark P, Pemberton DJ, Mareels G, Canuso CM, Fava M, Drevets WC. Efficacy and safety of aticaprant, a kappa receptor antagonist, adjunctive to oral SSRI/SNRI antidepressant in major depressive disorder: results of a phase 2 randomized, double-blind, placebo-controlled study. Neuropsychopharmacology. 2024;49(9):1437-47. doi:10.1038/s41386-024-01862-x. PMID:38649428.
29. ClinicalTrials.gov. A study of aticaprant as adjunctive therapy in adult participants with major depressive disorder with moderate-to-severe anhedonia (VENTURA-1). Identifier NCT05455684 [Internet]. Bethesda (MD): National Library of Medicine [cited 2026 May 13]. Available from: https://clinicaltrials.gov/study/NCT05455684
30. Johnson & Johnson. Johnson & Johnson statement on VENTURA program [Internet]. New Brunswick (NJ): Johnson & Johnson; 2025 Mar 6 [cited 2026 May 13]. Available from: https://www.jnj.com/media-center/press-releases/johnson-johnson-statement-on-ventura-program
31. Sumitomo Pharma Co Ltd, Otsuka Pharmaceutical Co Ltd. Sumitomo Pharma and Otsuka announce topline results from phase 3 DIAMOND 1 and DIAMOND 2 clinical studies evaluating ulotaront in schizophrenia [Internet]. Osaka and Tokyo: Sumitomo Pharma/Otsuka; 2023 Jul 31 [cited 2026 May 13]. Available from: https://www.sumitomo-pharma.com/news/20230731-1.html
32. ClinicalTrials.gov. Efficacy and safety of SEP-363856 in acutely psychotic people with schizophrenia. Identifier NCT04072354 [Internet]. Bethesda (MD): National Library of Medicine [cited 2026 May 13]. Available from: https://clinicaltrials.gov/study/NCT04072354
33. ClinicalTrials.gov. Efficacy and safety of SEP-363856 in acutely psychotic people with schizophrenia. Identifier NCT04092686 [Internet]. Bethesda (MD): National Library of Medicine [cited 2026 May 13]. Available from: https://clinicaltrials.gov/study/NCT04092686
34. van Dyck CH, Swanson CJ, Aisen P, Bateman RJ, Chen C, Gee M, et al. Lecanemab in early Alzheimer disease. N Engl J Med. 2023;388(1):9-21. doi:10.1056/NEJMoa2212948. PMID:36449413.
35. Budd Haeberlein S, Aisen PS, Barkhof F, Chalkias S, Chen T, Cohen S, et al. Two randomized phase 3 studies of aducanumab in early Alzheimer disease. J Prev Alzheimers Dis. 2022;9(2):197-210. doi:10.14283/jpad.2022.30.
36. Salloway S, Chalkias S, Barkhof F, Burkett P, Barakos J, Purcell D, Suhy J, et al. Amyloid-related imaging abnormalities in 2 phase 3 studies evaluating aducanumab in patients with early Alzheimer disease. JAMA Neurol. 2022;79(1):13-21. doi:10.1001/jamaneurol.

