Publication System Publication System

Regulatory Explainability in AI-Assisted Pharmaceutical Formulation and Process Development

Original Research | Open access | Published: 10 July 2024
Volume 0, article number 170, (0) Cite this article
You have full access to this open access article.
, ,
  1. Department of Pharmaceutical Technology and Drug Systems, Faculty of Pharmacy, Vietnam National University, Hanoi, Vietnam
  2. Department of Therapeutic Engineering and Applications, Faculty of Medicine, Can Tho University, Can Tho, Vietnam
110 Accesses

Abstract

Artificial intelligence is rapidly being adopted to optimise pharmaceutical formulation and manufacturing processes, yet its inherent opacity poses a fundamental challenge to regulatory frameworks built on transparency, scientific justification, and mechanistic understanding. In pharmaceutical development, AI systems may influence formulation selection, process parameter optimisation, release modelling, stability prediction, and scale-up strategy. These decisions can directly or indirectly affect product quality and patient safety. The regulatory issue is therefore not whether AI can improve development efficiency, but whether its outputs can be explained, justified, verified, and governed. There is currently no clear regulatory consensus on what constitutes sufficient explainability for AI models used in pharmaceutical formulation and process development. Existing expectations for pharmaceutical development presume that sponsors can describe the relationship between material attributes, process parameters, critical quality attributes, and clinical performance. Many AI models, particularly neural networks, ensemble methods, and adaptive systems, challenge this assumption because their internal decision logic may not be readily interpretable. This uncertainty creates difficulty for industry, regulators, and quality units seeking to evaluate whether AI-supported decisions are scientifically sound. This article identifies the regulatory gaps and risk domains associated with non-explainable AI in pharmaceutical development. It examines how AI is being used across formulation design, process development, manufacturing optimisation, and data-rich pharmaceutical quality systems. It then defines explainability and evidence requirements appropriate for different levels of regulatory risk. The central argument is that explainability should be treated as a regulatory quality attribute of AI systems, not merely as a technical preference. The proposed framework categorises AI applications according to their intended use, model complexity, degree of influence on quality decisions, and potential impact on patient safety. It links these categories to evidence requirements, including training data documentation, performance validation, uncertainty assessment, interpretability justification, human oversight, and lifecycle change controls. Four tables present the AI application landscape, regulatory gap analysis, explainability requirements, and the proposed framework. Together, these elements provide a structured basis for regulatory dialogue and future guidance development. A structured, risk-based approach to regulatory explainability can enable responsible adoption of AI while protecting the integrity of the pharmaceutical regulatory system. Low-risk AI tools may require documented performance and traceability, whereas high-risk tools that influence critical quality decisions require stronger interpretability, validation, and governance. The proposed approach does not require all models to be fully transparent, but it does require that the explanation provided be proportionate to the decision being supported. Regulatory explainability is therefore presented as a necessary bridge between innovation, quality assurance, and patient protection.

Explore related subjects
Discover the latest articles in related subjects:

Introduction

Artificial intelligence is moving from exploratory computational support into the core of pharmaceutical formulation and process development, where it is used to predict formulation behaviour, optimise experimental design, and support manufacturing decisions. Machine learning directed formulation development has been described as a way to accelerate product design by learning relationships between composition, processing conditions, and product performance [1]. This shift is particularly significant because formulation and process decisions shape the control strategy that later supports regulatory approval, commercial manufacturing, and lifecycle management. The promise of AI is therefore inseparable from the regulatory obligation to explain how development decisions are made.

The expanding literature on AI-enabled drug product development demonstrates that models can support solubility prediction, release profile modelling, stability assessment, nanocrystal design, 3D-printed medicines, and solid dosage form development [2]. These applications show that AI is not confined to discovery chemistry, but increasingly affects the pharmaceutical development space in which product quality is established [3]. At the same time, the technical performance of a predictive model does not automatically make it suitable for regulatory use. A model that predicts well but cannot be explained may create uncertainty when sponsors must justify critical material attributes, critical process parameters, and quality risk decisions.

The regulatory concern is sharpened by the wider debate on explainable artificial intelligence, especially in high-stakes settings where opaque models can obscure causal reasoning, accountability, and error detection. Rudin argues that black-box models should not be casually accepted for high-stakes decisions when interpretable alternatives are possible [4]. In healthcare, Ghassemi, Oakden-Rayner and Beam caution that some forms of post-hoc explanation may create false confidence rather than genuine understanding [5]. These concerns are directly relevant to pharmaceutical development, where a convincing visual or feature-importance explanation may not be enough to justify a formulation or process decision.

This article argues that regulatory explainability should be treated as a structured requirement for AI-assisted pharmaceutical formulation and process development. The purpose is to define the regulatory problem, contextualise AI applications, identify gaps in the current regulatory paradigm, and propose a risk-based framework for explainability and evidence. The framework is positioned as a regulatory science contribution, not as a replacement for existing pharmaceutical quality principles. It seeks to make AI-enabled innovation compatible with the expectations of transparency, reproducibility, lifecycle control, and patient protection.

Figure 1 illustrates the regulatory explainability gap that emerges when AI-supported formulation and process decisions influence product-quality reasoning without producing explanations that are scientifically inspectable, auditable, and acceptable for regulatory review.

Figure 1. Regulatory Explainability Gap in AI-Assisted Pharmaceutical Formulation and Process Development

Figure 1. Regulatory Explainability Gap in AI-Assisted Pharmaceutical Formulation and Process Development

Regulatory Problem

The central regulatory problem is that pharmaceutical quality frameworks assume that sponsors can explain how development knowledge supports the proposed control strategy. In conventional quality-by-design practice, scientific understanding links formulation variables and process parameters to product performance, even when empirical models are used. AI systems complicate this expectation because their predictions may emerge from high-dimensional correlations rather than explicit mechanistic reasoning [6]. When a model recommends a formulation composition or process setting without a transparent rationale, the sponsor may struggle to demonstrate that the decision is scientifically justified.

This problem is especially acute when AI models influence decisions that are later embedded in regulatory submissions, manufacturing controls, or post-approval change strategies. A landscape analysis of AI and machine learning in drug development regulatory submissions shows that regulators are already encountering these technologies across development programmes [7]. However, submission experience alone does not establish a consistent standard for what level of model documentation, explainability, or validation is sufficient. The result is a regulatory grey zone in which sponsors may innovate faster than assessment frameworks can adapt.

Regulatory review depends not only on evidence that a model performs well, but also on evidence that it is appropriate for its intended use. Explainable AI surveys emphasise that different explanation methods serve different purposes, including local interpretation, global model understanding, feature attribution, and surrogate modelling [8]. For pharmaceutical development, the regulatory question is not simply whether SHAP, LIME, attention maps, or surrogate models can be generated. The question is whether those explanations support a scientifically credible assessment of risk, control, and product quality.

Accountability becomes difficult when responsibility for AI-derived decisions is distributed across model developers, pharmaceutical scientists, data scientists, vendors, and quality units. In medicine, multidisciplinary analyses of explainability emphasise that technical explanations must be connected to the needs of users, institutions, and governance systems [9]. The same logic applies to pharmaceutical development because responsibility cannot be transferred from the manufacturer to an algorithm. If an AI model contributes to a poor formulation choice, unstable product, or unsuitable process parameter, the regulatory system still expects the sponsor to explain, control, and correct the decision.

AI-Assisted Formulation and Process Development Context

AI-assisted formulation development includes the use of machine learning to predict how excipient combinations, drug properties, processing conditions, and dosage form design influence critical performance attributes. Bannigan and colleagues describe machine learning as a tool that can guide drug formulation development by learning from prior experimental data and narrowing the formulation design space [1]. Related reviews highlight applications in solid oral dosage forms, including disintegration, dissolution, manufacturability, and tablet quality prediction [10]. These uses are attractive because formulation development is data-rich, experimentally expensive, and often constrained by nonlinear interactions among materials and processes.

In process development, AI methods are increasingly aligned with data-intensive manufacturing strategies, including process analytical technology, continuous manufacturing, and real-time process optimisation. Reviews of artificial neural networks in pharmaceutical formulation show their use in predicting, characterising, and optimising formulation and process outcomes [3]. Pharma 4.0 discussions also position AI as part of a broader transformation involving digital manufacturing, automation, advanced analytics, and cyber-physical production systems [11]. Regulatory explainability becomes essential when these models influence process design, scale-up, or control strategy decisions that affect commercial product quality.

The model landscape is diverse, ranging from interpretable regression and decision trees to ensemble methods, artificial neural networks, deep learning systems, and hybrid approaches that combine mechanistic knowledge with data-driven prediction. In vitro formulation development has already been examined through interpretable machine learning methods, illustrating that explainability can be built into formulation modelling rather than treated only as an afterthought [12]. For nanocrystals, AI has been used to predict formulation feasibility and product characteristics, demonstrating both the potential and the complexity of model-driven formulation decisions [13]. Table 1 categorises AI applications in pharmaceutical formulation and process development and their typical model types.

Table 1. AI Applications in Pharmaceutical Formulation and Process Development: Use Cases, Data Types, and Model Architectures

AI application area

Representative pharmaceutical use case

Common data types

Typical model architectures

Regulatory explainability concern

Formulation screening

Selection of excipients, drug loading, and composition ranges

Experimental formulation datasets, physicochemical descriptors, excipient properties

Regression models, random forests, gradient boosting, neural networks

Whether the model rationale supports a scientifically justified design space

Solubility and dissolution prediction

Prediction of solubility enhancement, dissolution rate, and release behaviour

Molecular descriptors, formulation composition, dissolution profiles

Support vector machines, ensemble models, artificial neural networks

Whether feature contributions align with pharmaceutical understanding

Stability prediction

Forecasting degradation, shelf-life risk, and formulation robustness

Stability time points, stress conditions, packaging variables, impurity trends

Neural networks, deep learning, time-series models

Whether predictions can be justified for storage conditions and specification setting

Nanocrystal and particulate systems

Prediction of particle size, stabiliser choice, and nanosuspension feasibility

Drug properties, stabiliser descriptors, process parameters

Classification models, ensemble learning, neural networks

Whether model outputs can be linked to colloidal stability and manufacturability

3D-printed medicines

Prediction of printability, dose geometry, release behaviour, and manufacturing feasibility

Printing parameters, polymer properties, geometry, mechanical data

Deep learning, classification models, hybrid predictive models

Whether the AI explanation supports dose accuracy and patient-specific manufacturing controls

Solid dosage form development

Tablet compression, disintegration, dissolution, and manufacturability optimisation

Powder properties, compression data, process settings, tablet attributes

Artificial neural networks, random forests, Gaussian processes

Whether critical material and process attributes remain traceable

Process optimisation

Identification of operating ranges for blending, granulation, drying, and compression

Process analytical technology signals, batch records, sensor data

Multivariate models, reinforcement learning, ensemble methods

Whether recommended process changes can be validated and controlled

Scale-up support

Translation from laboratory or pilot scale to commercial manufacturing

Scale-dependent process data, equipment parameters, quality attributes

Hybrid models, Bayesian models, surrogate models

Whether scale-up decisions are explainable enough for regulatory review and lifecycle management

AI is also changing adjacent areas of pharmaceutical development, including 3D printing, personalised dosage forms, and platform-based digital formulation tools. Machine learning has been used to predict printability of medicines, while broader work on AI in drug delivery design shows the increasing role of computational decision support in dosage form innovation [14]. Web-based formulation platforms further suggest that AI may become embedded in routine development workflows rather than reserved for specialised projects [15]. As these tools become more operationally integrated, the need for explainable, auditable, and submission-ready evidence becomes a regulatory priority rather than a theoretical concern.

Current Regulatory Gap

The current regulatory gap is not the absence of pharmaceutical quality principles, but the absence of AI-specific interpretation of those principles for formulation and process development. Existing regulatory expectations are grounded in development knowledge, validation, traceability, and control, yet they do not define how an AI model should be explained when it influences formulation choice, process optimisation, or scale-up. Reviews of AI-enabled pharmaceutical formulation show that these systems are already being used to make decisions that resemble development rationales traditionally justified through experiments and mechanistic reasoning [16]. This creates a mismatch between the speed of AI adoption and the specificity of regulatory expectations.

Current quality-by-design concepts can accommodate modelling, but they were not designed for continuously updated, high-dimensional, and partially opaque AI systems. Artificial neural networks and deep learning models used in formulation and clinical-outcome-related quality-by-design applications can identify complex patterns, but their internal representations may not correspond to familiar pharmaceutical mechanisms [17]. A regulatory submission based on such models would need more than accuracy metrics; it would need a defensible explanation of how the model supports product understanding. Without that explanation, AI risks becoming a hidden layer between experimental evidence and regulatory judgement.

A second gap concerns documentation standards for training data, feature engineering, model selection, model versioning, and model maintenance. Emerging pharmaceutical AI platforms demonstrate how model-driven formulation design can become scalable and user-facing, but standardised expectations for documenting the underlying data provenance and algorithmic assumptions remain underdeveloped [15]. Table 2 summarises the mismatch between current regulatory expectations and AI-specific challenges.

Table 2. Regulatory Gap Analysis: How Current ICH and GMP Frameworks Fall Short for AI-Assisted Development

Current regulatory expectation

Conventional interpretation in pharmaceutical development

AI-specific challenge

Consequence for regulatory review

Proposed direction

Scientific understanding

Sponsor explains how formulation and process variables affect product quality

Black-box models may predict outcomes without mechanistic transparency

Reviewers may be unable to judge whether the model supports a credible control strategy

Require risk-based explainability linked to intended use

Model validation

Models are justified through fit, prediction, and confirmatory evidence

Train/test performance may not reflect future batches, new materials, or scale-up conditions

Apparent performance may not translate into regulated manufacturing settings

Require prospective and lifecycle validation for higher-risk uses

Data integrity

Data are attributable, legible, contemporaneous, original, and accurate

Training datasets may combine historical, public, vendor, and synthetic data

Bias, missingness, and undocumented preprocessing may remain hidden

Require data provenance, curation records, and dataset relevance assessment

Change control

Process and method changes are evaluated before implementation

Adaptive AI models may change after deployment or retraining

Model drift may affect quality decisions without formal regulatory awareness

Require predetermined model change plans and revalidation triggers

Auditability

Decisions can be reconstructed from records and procedures

AI outputs may be difficult to reproduce if code, parameters, or data versions change

Regulatory inspection may not reconstruct the basis for a decision

Require model version control, audit trails, and reproducible inference records

Human oversight

Qualified personnel review and approve quality decisions

Automation may encourage overreliance on model recommendations

Accountability may become blurred across scientists, vendors, and algorithms

Require defined human decision authority and escalation pathways

Lifecycle management

Development knowledge supports post-approval changes and continuous improvement

AI may generate new recommendations as data accumulate

Post-approval changes may be difficult to justify if model logic is unclear

Require lifecycle evidence packages and explainability maintenance

The third gap lies in the difference between technical validation and regulatory validation. In machine learning, model performance may be established through retrospective train/test splitting, cross-validation, or external test sets, but pharmaceutical regulation often requires prospective evidence that a proposed control strategy performs under defined operating conditions. This distinction is important because AI in drug development has expanded from discovery into development workflows, where errors can affect regulatory dossiers and patient-facing products [18]. A regulatory framework must therefore translate technical validation into evidence that is meaningful for quality assessors, inspectors, and lifecycle reviewers.

Risk Domains

The first risk domain is patient safety, because formulation or process decisions that appear technical can ultimately alter dose delivery, exposure, stability, and therapeutic performance. AI tools used in drug delivery design may optimise complex formulation variables, but incorrect predictions could lead to unsuitable excipient levels, poor release behaviour, or unstable products [19]. In 3D-printed medicines, for example, AI-supported printability or release predictions may influence dose geometry and patient-specific product performance [20]. If such predictions are not explainable, errors may be detected late, after development resources have been committed or after unsuitable assumptions have shaped the control strategy.

The second risk domain is product quality, especially where AI models learn from biased, sparse, or non-representative datasets. Pharmaceutical development datasets often reflect historical project choices, failed experiments that were not fully recorded, supplier-specific materials, and scale-dependent process behaviour. Machine learning can amplify these limitations when feature importance, training boundaries, and uncertainty are not transparent [21]. A model that performs well within a narrow historical design space may fail when applied to a new active ingredient, dosage form, equipment configuration, or manufacturing site.

The third risk domain is regulatory non-compliance, including delayed approvals, rejected modelling arguments, inadequate responses to regulatory questions, and weak post-approval change justifications. Regulatory agencies are already seeing AI and machine learning in drug development submissions, but the evidentiary expectations remain heterogeneous across contexts [7]. If a sponsor cannot explain how a model was trained, validated, governed, and linked to product quality, the model may be treated as exploratory rather than decision-supporting evidence. This risk is particularly serious when AI outputs are used to justify design space boundaries, critical process parameter ranges, or reduced experimental testing.

The fourth risk domain is business and lifecycle risk, because non-explainable AI can become difficult to maintain, transfer, or defend after approval. AI-supported medicinal chemistry and development workflows show that advanced algorithms can accelerate decision-making, but the long-term value of those decisions depends on reproducibility and institutional memory [22]. In manufacturing, a model that cannot be explained may obstruct technology transfer, site changes, raw material changes, or continuous improvement. Regulatory explainability therefore protects not only patients and regulators, but also the sponsor’s ability to manage product knowledge over the lifecycle.

Explainability and Evidence Requirements

Explainability should be understood as a spectrum rather than a binary property. At one end are inherently interpretable models, such as linear models, decision trees, sparse rule-based systems, and constrained models whose structure can be directly inspected. At the other end are black-box models, including deep neural networks and complex ensembles, where post-hoc methods may be needed to approximate the reasons for a prediction [3]. For pharmaceutical regulation, the key issue is whether the explanation is sufficiently faithful, stable, and relevant to the quality decision being supported.

Different explanation methods provide different kinds of evidence, and they should not be treated as interchangeable. Feature-attribution methods may identify variables associated with a prediction, surrogate models may approximate local behaviour, and attention-based explanations may suggest which inputs influenced a model, but none of these automatically proves causal understanding [8]. Medical AI scholarship has warned that explanations can create an illusion of understanding if they are not validated against domain knowledge and user needs [5]. In pharmaceutical development, explanation methods must therefore be assessed against formulation science, process understanding, and the intended regulatory use of the model.

A risk-based approach should link the required level of explainability to the impact of the AI-supported decision. Low-risk tools used for literature triage, exploratory screening, or hypothesis generation may require data documentation, model performance metrics, and user limitations, but not full mechanistic transparency. Moderate-risk tools used to prioritise experiments, predict stability, or support formulation selection should provide interpretable features, uncertainty estimates, and confirmatory experimental evidence [23]. High-risk tools that directly influence critical quality attributes, release specifications, design space claims, or commercial control strategies require robust explanation, prospective validation, and human decision accountability.

The concept of causability is particularly useful because it focuses on whether an explanation is usable by human experts in a specific decision context. Holzinger and colleagues distinguish explainability from the practical ability of users to understand and act on AI outputs, which is essential in medical and regulated settings [24]. For pharmaceutical development, an explanation should allow formulation scientists, process engineers, quality units, and regulators to evaluate whether the model’s reasoning is plausible. Table 3 maps risk-based explainability requirements to AI model complexity and intended use.

Table 3. Risk-Based Explainability and Evidence Requirements for AI-Assisted Pharmaceutical Applications

Risk tier

Intended AI use

Typical model complexity

Minimum explainability expectation

Required evidence package

Regulatory acceptance condition

Tier 1: Exploratory support

Literature triage, formulation hypothesis generation, early screening prioritisation

Low to high, including black-box models

Clear statement of intended use, limitations, and non-decisional status

Dataset description, performance summary, user instructions, bias screening

Acceptable if outputs are not used as primary evidence for quality decisions

Tier 2: Development decision support

Selection of formulation candidates, experiment prioritisation, process parameter exploration

Interpretable models, ensembles, neural networks with post-hoc explanations

Feature-level explanation, uncertainty estimate, applicability domain, expert review

Training data provenance, validation dataset, explanation report, confirmatory experiments

Acceptable if AI supports but does not replace scientific justification

Tier 3: Quality-relevant modelling

Stability prediction, dissolution modelling, scale-up support, design space proposal

Moderate to high complexity, including hybrid and surrogate models

Validated explanation linked to critical quality attributes and process understanding

Prospective validation, sensitivity analysis, error impact assessment, model version control

Acceptable if explanations are scientifically plausible and experimentally confirmed

Tier 4: Control strategy influence

Real-time process optimisation, release-related decisions, automated process adjustments

High complexity models, adaptive models, reinforcement learning, digital twins

High-fidelity interpretability, predefined operating boundaries, human override logic

Lifecycle validation, audit trails, drift monitoring, change control plan, governance record

Acceptable only where model behaviour is auditable, bounded, and continuously governed

Tier 5: Non-acceptable autonomous use

Autonomous decisions affecting patient safety without human review or explainable rationale

Opaque or continuously changing systems

No sufficient explanation for regulatory accountability

Evidence insufficient by design

Not acceptable for critical quality or safety decisions without redesign

Evidence requirements should also include negative evidence, meaning documentation of model limitations, failure modes, and conditions under which the model should not be used. Surveys of explainable AI for medical applications show that explanation reliability depends on the model, the user, and the decision context [21]. In pharmaceutical development, this means sponsors should document the applicability domain, out-of-distribution behaviour, dataset gaps, and known sources of bias. A regulatory-quality explanation should make uncertainty visible rather than conceal it behind a persuasive prediction.

Proposed Regulatory Framework

The proposed framework treats explainability as a risk-based regulatory requirement embedded within pharmaceutical development, not as a universal demand for complete algorithmic transparency. Low-risk AI applications should be permitted with documented intended use, dataset provenance, basic performance validation, and clear limits on decision authority. Moderate-risk applications should require validated explanation methods, expert review, uncertainty assessment, and confirmatory experiments. High-risk applications should require interpretable or explainable-by-design models unless the sponsor can demonstrate that a black-box system is bounded, auditable, prospectively validated, and subject to strict human oversight.

The first element of the framework is an intended-use statement that defines exactly how the AI output will influence development or manufacturing decisions. This mirrors the regulatory logic used in digital health and AI-based medical device oversight, where the risk of an algorithm depends on its role in clinical or operational decisions [25]. In pharmaceutical development, the intended-use statement should distinguish exploratory use from quality-relevant use. A model used to suggest candidate excipients is different from a model used to justify a commercial design space or approve a real-time manufacturing adjustment.

The second element is an AI model evidence package that includes data provenance, preprocessing rationale, model architecture, training strategy, validation design, explanation method, uncertainty characterisation, and governance controls. Reviews of AI in drug development emphasise the breadth of applications and the need to align model performance with development purpose [26]. For pharmaceutical development, this evidence package should be written in a way that both technical reviewers and quality assessors can understand. It should also preserve enough detail for inspection, reproducibility, lifecycle management, and post-approval change evaluation.

The third element is a predetermined model change control plan for retraining, updating, or replacing AI models used in regulated development or manufacturing contexts. Comparative analyses of AI and machine-learning-based medical device approvals show that adaptive algorithms create distinctive regulatory challenges when model behaviour can evolve after initial review [27]. Pharmaceutical applications raise similar questions when new batch data, supplier data, equipment data, or stability data are added to a model. Table 4 presents the proposed regulatory framework and its risk-tiered elements.

Table 4. Proposed Regulatory Framework for Explainable AI in Pharmaceutical Formulation and Process Development

Framework component

Low-risk AI application

Moderate-risk AI application

High-risk AI application

Regulatory purpose

Intended-use classification

Exploratory or advisory use only

Supports development decisions with human review

Influences critical quality, control strategy, or lifecycle decisions

Defines the regulatory weight of AI output

Data governance

Basic dataset description and source traceability

Full data provenance, preprocessing record, representativeness assessment

Data integrity controls, audit trail, dataset lock, bias and drift monitoring

Ensures that model evidence is attributable and reviewable

Model documentation

Model type, version, performance metrics, limitations

Architecture, feature engineering, validation design, explanation method

Complete model dossier, reproducible inference record, source/version control

Makes the AI system inspectable and reproducible

Explainability requirement

User-facing rationale and limitations

Feature attribution, surrogate explanation, sensitivity analysis, expert review

Explainable-by-design model or validated high-fidelity explanation

Links model logic to scientific and quality understanding

Validation expectation

Retrospective or exploratory validation

External validation and confirmatory experiments

Prospective validation under intended operating conditions

Aligns technical validation with regulatory confidence

Human oversight

Scientist reviews outputs before use

Qualified expert approves decision relevance

Quality unit and accountable decision-maker approve use and escalation

Prevents transfer of responsibility to the algorithm

Change control

Version record for major changes

Predefined retraining triggers and revalidation criteria

Predetermined change control plan with regulatory impact assessment

Supports lifecycle management and post-approval control

Regulatory outcome

Acceptable as non-primary evidence

Acceptable as supportive evidence with justification

Acceptable only with strong explainability, validation, and governance

Protects innovation while preserving product quality and patient safety

The fourth element is a proportional acceptance rule that prevents both overregulation and underregulation. AI should not be excluded merely because it is complex, but complexity should raise the evidentiary burden when the model influences product quality or patient safety. Reviews of AI in pharmaceutical technology and drug delivery show that increasingly sophisticated systems are likely to become routine in development workflows [28]. A tiered framework allows regulators to encourage innovation while requiring stronger explanation where the consequences of error are greater.

Figure 2 presents the proposed risk-based framework in which explainability and evidence requirements increase as AI outputs move from exploratory support toward quality-relevant modelling, control-strategy influence, and patient-safety-sensitive decisions.

Figure 2. Risk-Based Explainability and Evidence Framework for AI-Supported Pharmaceutical Decisions

Figure 2. Risk-Based Explainability and Evidence Framework for AI-Supported Pharmaceutical Decisions

Implementation Pathway

Implementation should begin with harmonised guidance that translates existing pharmaceutical quality expectations into AI-specific requirements. Such guidance should define categories of AI use, minimum documentation standards, acceptable validation strategies, and risk-based explainability expectations. The literature on AI in drug discovery and development shows rapid methodological expansion, making fragmented national expectations inefficient for global development programmes [29]. An ICH-style approach would help sponsors build evidence packages that are acceptable across regions.

A second implementation step is pre-competitive collaboration on model documentation templates, benchmark datasets, and explanation validation practices. Pharmaceutical AI spans formulation development, drug delivery, manufacturing analytics, and broader development decision-making, so common documentation standards would reduce uncertainty for both sponsors and regulators [30]. These standards should not prescribe a single algorithm or explanation method. Instead, they should define what a sponsor must be able to demonstrate when AI is used to support a regulated development decision.

A third implementation step is workforce development for both regulators and industry. Regulatory assessors need sufficient familiarity with machine learning concepts, explainability methods, validation pitfalls, and data governance to evaluate AI evidence critically. Industry scientists need to understand that model performance is not the same as regulatory acceptability, particularly when black-box systems are used in high-stakes quality contexts [4]. Cross-functional training should involve formulation scientists, process engineers, statisticians, data scientists, quality professionals, and regulatory affairs specialists.

A fourth implementation step is the use of pilot programmes and controlled regulatory experiments. Machine learning for small-molecule drug discovery has already moved across academic and industrial settings, showing that translational governance is as important as technical performance [31]. Pilot programmes could allow sponsors to submit AI model evidence packages for scientific advice before using them in pivotal development or post-approval contexts. This would create shared learning while avoiding premature reliance on opaque systems in critical regulatory decisions.

Figure 3 outlines the implementation pathway through which regulatory explainability can move from general principle to operational practice through harmonised guidance, model-documentation standards, cross-functional capability building, pilot submissions, lifecycle governance, and regulatory convergence.

Figure 3. Implementation Pathway for Regulatory Explainability in AI-Assisted Pharmaceutical Development

Figure 3. Implementation Pathway for Regulatory Explainability in AI-Assisted Pharmaceutical Development

Accountability and Limitations

The manufacturer must remain accountable for AI-derived decisions, regardless of whether the model was developed internally, licensed from a vendor, or embedded in a software platform. Explainability in healthcare has been described as a sociotechnical requirement involving users, institutions, and governance, not merely a technical output [9]. In pharmaceutical development, this means the sponsor must retain decision authority, maintain documentation, review model outputs, and justify how AI evidence supports product quality. Human oversight should therefore be formalised as part of the quality system rather than treated as informal scientific judgement.

The proposed framework has limitations because explainability methods remain technically imperfect and context-dependent. Post-hoc explanations may be unstable, incomplete, or misleading, especially for complex models whose internal logic does not correspond to causal mechanisms [5]. Conversely, insisting on fully interpretable models in every case could unnecessarily restrict useful innovation. The framework therefore adopts proportionality, but proportionality itself requires regulatory judgement and continued refinement as methods evolve.

International convergence is another limitation and implementation challenge. AI-based medical device regulation already shows variation between jurisdictions, and similar divergence could emerge in pharmaceutical AI if regulators do not coordinate expectations [27]. The framework also depends on data quality, organisational maturity, and the availability of experts who can evaluate both pharmaceutical science and AI methodology. For this reason, regulatory explainability should be developed as an evolving discipline within regulatory science, not as a one-time checklist.

Conclusion

AI-assisted formulation and process development is becoming an important part of pharmaceutical innovation, but its regulatory value depends on whether AI-supported decisions can be explained, justified, validated, and governed. The central challenge is not simply that AI models are complex, but that their complexity can obscure the scientific reasoning needed to protect product quality and patient safety. Regulatory explainability must therefore become a formal expectation for AI systems that influence pharmaceutical development decisions.

A risk-based framework provides a practical path between uncritical adoption and excessive restriction. It allows low-risk exploratory tools to be used with proportionate documentation, while requiring stronger evidence, interpretability, validation, and governance when AI affects critical quality decisions. This approach preserves the benefits of AI while maintaining the regulatory principles of transparency, accountability, and lifecycle control.

The next step is coordinated action among regulators, industry, academia, and standards-setting bodies. Harmonised guidance, pilot programmes, model documentation standards, and workforce training are needed to bring regulatory science into the AI era. Explainable AI should not be treated as an optional technical feature, but as a condition for trustworthy pharmaceutical development.

Acknowledgements

None

Conflict of interest

None

Financial support

None

Ethics statement

None

References

Bannigan P, Aldeghi M, Bao Z, Häse F, Aspuru-Guzik A, Allen C. Machine learning directed drug formulation development. Adv Drug Deliv Rev. 2021;175:113806.
Murray JD, Lange JJ, Bennett-Lenane H, Holm R, Kuentz M, O'Dwyer PJ, et al. Advancing algorithmic drug product development: Recommendations for machine learning approaches in drug formulation. Eur J Pharm Sci. 2023;191:106562.
Wang S, Di J, Wang D, Dai X, Hua Y, Gao X, et al. State-of-the-art review of artificial neural networks to predict, characterize and optimize pharmaceutical formulation. Pharmaceutics. 2022;14(1):183.
Rudin C. Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead. Nat Mach Intell. 2019;1(5):206-15.
Ghassemi M, Oakden-Rayner L, Beam AL. The false hope of current approaches to explainable artificial intelligence in health care. Lancet Digit Health. 2021;3(11):e745-50.
Guidotti R, Monreale A, Ruggieri S, Turini F, Giannotti F, Pedreschi D. A survey of methods for explaining black box models. ACM Comput Surv. 2018;51(5):93.
Liu Q, Huang R, Hsieh J, Zhu H, Tiwari M, Liu G, et al. Landscape analysis of the application of artificial intelligence and machine learning in regulatory submissions for drug development from 2016 to 2021. Clin Pharmacol Ther. 2023;113(4):771-4.
Arrieta AB, Díaz-Rodríguez N, Del Ser J, Bennetot A, Tabik S, Barbado A, et al. Explainable Artificial Intelligence (XAI): Concepts, taxonomies, opportunities and challenges toward responsible AI. Inf Fusion. 2020;58:82-115.
Amann J, Blasimme A, Vayena E, Frey D, Madai VI; Precise4Q Consortium. Explainability for artificial intelligence in healthcare: A multidisciplinary perspective. BMC Med Inform Decis Mak. 2020;20(1):310.
Lou H, Lian B, Hageman MJ. Applications of machine learning in solid oral dosage form development. J Pharm Sci. 2021;110(9):3150-65.
Malheiro V, Duarte J, Veiga F, Mascarenhas-Melo F. Exploiting Pharma 4.0 technologies in the non-biological complex drugs manufacturing: Innovations and implications. Pharmaceutics. 2023;15(11):2545.
Ye Z, Yang W, Yang Y, Ouyang D. Interpretable machine learning methods for in vitro pharmaceutical formulation development. Food Front. 2021;2(2):195-207.
He Y, Ye Z, Liu X, Wei Z, Qiu F, Li HF, et al. Can machine learning predict drug nanocrystals? J Control Release. 2020;322:274-85.
Elbadawi M, Castro BM, Gavins FK, Ong JJ, Gaisford S, Pérez G, et al. M3DISEEN: A novel machine learning approach for predicting the 3D printability of medicines. Int J Pharm. 2020;590:119837.
Elbadawi M, McCoubrey LE, Gavins FK, Ong JJ, Goyanes A, Gaisford S, et al. Disrupting 3D printing of medicines with machine learning. Trends Pharmacol Sci. 2021;42(9):745-57.
Bao Z, Bufton J, Hickman RJ, Aspuru-Guzik A, Bannigan P, Allen C. Revolutionizing drug formulation development: The increasing impact of machine learning. Adv Drug Deliv Rev. 2023;202:115108.
Simões MF, Silva G, Pinto AC, Fonseca M, Silva NE, Pinto RM, et al. Artificial neural networks applied to quality-by-design: From formulation development to clinical outcome. Eur J Pharm Biopharm. 2020;152:282-95.
Vamathevan J, Clark D, Czodrowski P, Dunham I, Ferran E, Lee G, et al. Applications of machine learning in drug discovery and development. Nat Rev Drug Discov. 2019;18(6):463-77.
Hassanzadeh P, Atyabi F, Dinarvand R. The significance of artificial intelligence in drug delivery system design. Adv Drug Deliv Rev. 2019;151-152:169-90.
Elbadawi M, McCoubrey LE, Gavins FK, Ong JJ, Goyanes A, Gaisford S, et al. Harnessing artificial intelligence for the next generation of 3D printed medicines. Adv Drug Deliv Rev. 2021;175:113805.
Tjoa E, Guan C. A survey on explainable artificial intelligence (XAI): Toward medical XAI. IEEE Trans Neural Netw Learn Syst. 2021;32(11):4793-813.
Struble TJ, Alvarez JC, Brown SP, Chytil M, Cisar J, DesJarlais RL, et al. Current and future roles of artificial intelligence in medicinal chemistry synthesis. J Med Chem. 2020;63(16):8667-82.
Ajdarić J, Ibrić S, Pavlović A, Ignjatović L, Ivković B. Prediction of drug stability using deep learning approach: Case study of esomeprazole 40 mg freeze-dried powder for solution. Pharmaceutics. 2021;13(6):829.
Holzinger A, Langs G, Denk H, Zatloukal K, Müller H. Causability and explainability of artificial intelligence in medicine. Wiley Interdiscip Rev Data Min Knowl Discov. 2019;9(4):e1312.
Benjamens S, Dhunnoo P, Meskó B. The state of artificial intelligence-based FDA-approved medical devices and algorithms: An online database. NPJ Digit Med. 2020;3(1):118.
Mak KK, Pichika MR. Artificial intelligence in drug development: Present status and future prospects. Drug Discov Today. 2019;24(3):773-80.
Muehlematter UJ, Daniore P, Vokinger KN. Approval of artificial intelligence and machine learning-based medical devices in the USA and Europe (2015–20): A comparative analysis. Lancet Digit Health. 2021;3(3):e195-203.
Vora LK, Gholap AD, Jetha K, Thakur RR, Solanki HK, Chavda VP. Artificial intelligence in pharmaceutical technology and drug delivery design. Pharmaceutics. 2023;15(7):1916.
Mak KK, Wong YH, Pichika MR. Artificial intelligence in drug discovery and development. In: Hock FJ, Gralinski MR, editors. Drug Discovery and Evaluation: Safety and Pharmacokinetic Assays. Cham: Springer; 2024. p. 1461-98.
Patel V, Shah M. Artificial intelligence and machine learning in drug discovery and development. Intell Med. 2022;2(3):134-40.
Volkamer A, Riniker S, Nittinger E, Lanini J, Grisoni F, Evertsson E, et al. Machine learning for small molecule drug discovery in academia and industry. Artif Intell Life Sci. 2023;3:100056.

Author information

Nguyen Thanh Huy, Pham Quang Minh & Le Thi Bich contributed to this work.

Authors and affiliations

Department of Pharmaceutical Technology and Drug Systems, Faculty of Pharmacy, Vietnam National University, Hanoi, Vietnam
Nguyen Thanh Huy & Pham Quang Minh

Department of Therapeutic Engineering and Applications, Faculty of Medicine, Can Tho University, Can Tho, Vietnam
Le Thi Bich

Corresponding author

Correspondence to Nguyen Thanh Huy

Rights and permissions

Open Access The author(s) retain copyright. This article is licensed under the Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License. It may be shared and adapted for non-commercial purposes with appropriate attribution, an indication of changes, and distribution of adaptations under the same license. Third-party material may be subject to separate terms identified in its credit line. View the license at https://creativecommons.org/licenses/by-nc-sa/4.0/.

About this article

Cite this article

Vancouver
Huy NT, Minh PQ, Bich LT. Regulatory Explainability in AI-Assisted Pharmaceutical Formulation and Process Development. . 0;0:170.
APA
Huy, N. T., Minh, P. Q., & Bich, L. T. (0). Regulatory Explainability in AI-Assisted Pharmaceutical Formulation and Process Development. EAMD 3, 0, 170.
Received
02 April 2024
Revised
06 May 2024
Accepted
01 June 2024
Published
10 July 2024
Version of record
10 July 2024

Share this article

Easily share this article with others using the link below:

Regulatory Explainability in AI-Assisted Pharmaceutical Formulation and Process Development
Scan to access
this article

Ready to submit?
Start a new submission or continue a submission in progress:
Submission Portal Author Guidelines

Follow this journal
Get notified of new updates and articles.