Publication System Publication System

Explainable AI-Generated Formulation Designs for Regulatory and Scientific Decision-Making

Original Research | Open access | Published: 10 January 2026
Volume 0, article number 192, (0) Cite this article
You have full access to this open access article.
,
  1. Department of Pharmaceutical Innovation and Systems, Faculty of Pharmacy, National and Kapodistrian University of Athens, Athens, Greece
104 Accesses

Abstract

Generative artificial intelligence is becoming increasingly relevant to pharmaceutical formulation because it can propose compositions, excipient combinations, processing conditions, and optimisation trajectories that may not be obvious through conventional experimental design. These capabilities create the possibility of faster development, broader exploration of formulation space, and more systematic use of prior knowledge. Yet the same models that expand formulation creativity often operate through complex latent representations that are difficult to interpret. This creates a trust problem for both scientific and regulatory decision-making. The central problem is that an AI-generated formulation is not only a predicted technical solution but also a claim about product performance, manufacturability, and quality. If the rationale behind that claim cannot be explained, formulation scientists may struggle to convert model outputs into mechanistic understanding. Regulators may likewise find it difficult to assess whether the proposed formulation is supported by transparent evidence. Opaque formulation design therefore risks becoming a translational bottleneck rather than an innovation accelerator. This perspective develops a conceptual framework for dual-purpose explainability in AI-generated pharmaceutical formulation design. The framework is designed to serve two decision contexts simultaneously. Scientific decision-makers require explanations that clarify formulation logic, reveal influential variables, and support hypothesis generation. Regulatory decision-makers require explanations that are auditable, reproducible, uncertainty-aware, and connected to product quality and safety evidence. The article first defines the conceptual gap between existing AI formulation capabilities and explainability expectations. It then describes the logic of AI-generated formulation, identifies distinct scientific and regulatory explanation requirements, and analyses transparency barriers. The proposed framework integrates global model explanations, local formulation-specific explanations, mechanistic interpretation, uncertainty communication, and regulatory evidence packaging. Three tables summarise the gap analysis, explainability requirements, and framework architecture. The article concludes that explainability must be treated as a design requirement rather than a post hoc add-on to pharmaceutical AI. AI-generated formulation designs will become useful only when their rationale can be interrogated, documented, challenged, and connected to established principles of product and process understanding. A dual-purpose explainability framework can help move the field from black-box prediction toward transparent, accountable, and scientifically meaningful formulation intelligence.

Explore related subjects
Discover the latest articles in related subjects:

Introduction

Artificial intelligence has moved from being a peripheral computational aid to becoming a central design instrument in drug discovery and pharmaceutical development. Early demonstrations of automated chemical design showed that machine learning could generate candidate structures through learned representations rather than through manually enumerated design rules [1]. Generative and reinforcement-learning systems extended this logic by proposing molecular candidates within high-dimensional design spaces, suggesting that algorithmic generation could become a general strategy for exploring complex pharmaceutical possibilities [2, 3]. In formulation science, the same generative logic is now increasingly relevant because formulation design also involves large, constrained, and interdependent design spaces.

The promise of AI-generated formulation design lies in its ability to connect prior formulation data, physicochemical descriptors, excipient properties, processing variables, and performance outcomes into predictive design systems. Reviews of deep learning in drug discovery show that these models can extract complex patterns from heterogeneous data, while machine learning-directed formulation development has demonstrated the value of algorithmic search for composition and process selection [4, 5]. In this setting, formulation design is no longer limited to incremental screening around familiar excipient combinations. It can become a computationally expanded search for feasible product architectures that satisfy performance, stability, manufacturability, and patient-use constraints.

However, the movement from prediction to decision creates an explanation problem. General surveys of explainable artificial intelligence emphasise that black-box models require interpretive tools when they are used in consequential settings, but pharmaceutical formulation imposes domain-specific explanation demands that exceed generic feature-attribution outputs [6, 7]. Formulation scientists do not only ask whether a model predicts dissolution, stability, printability, or viscosity. They ask why a particular drug–excipient–process relationship is plausible, whether the explanation is consistent with known mechanisms, and whether the proposed design can guide further experimentation.

This article argues that explainability must become a core requirement for AI-generated pharmaceutical formulations, particularly when such outputs may influence scientific development programmes or regulatory submissions. Regulatory-facing discussions of artificial intelligence in health care have already shown that model transparency, validation, and traceability are central to institutional trust, especially when algorithmic decisions can affect patient safety or product quality [8, 9]. The objective of this perspective is therefore to propose a dual-purpose explainability framework that serves formulation scientists and regulatory assessors at the same time. The framework is intended to convert AI-generated formulation designs from opaque suggestions into interpretable, auditable, and decision-ready scientific claims.

Conceptual Gap

The first conceptual gap is that current AI formulation systems are often evaluated mainly by predictive performance, search efficiency, or experimental acceleration rather than by the explanatory adequacy of their outputs. Machine learning-directed formulation studies have shown that models can support drug product development, biologics formulation, 3D-printability prediction, suspension design, and vaccine formulation optimisation [10-14]. Yet these achievements do not automatically answer why a proposed composition is scientifically reasonable or how the underlying rationale should be represented in a regulatory evidence package. A high-performing model can still fail as a decision tool if its reasoning cannot be scrutinised.

The second gap is that existing XAI methods were largely developed as general-purpose tools, not as formulation-specific instruments. Surveys of XAI describe feature importance, local surrogate models, rule extraction, attention-based interpretation, and counterfactual explanations as broad methodological families [6, 7]. These tools can be useful, but formulation design requires explanations that connect variables to formulation mechanisms such as solubilisation, solid-state stabilisation, release modulation, excipient compatibility, process-induced transformation, and manufacturability. An explanation that identifies an influential variable without clarifying its formulation role may be technically transparent but scientifically incomplete.

The third gap concerns the difference between scientific and regulatory decision-making. Scientists need explanations that help them generate hypotheses, reject implausible model proposals, and design follow-up experiments; human-centred XAI research has shown that trust depends on whether explanations support the user’s actual reasoning task [15]. Regulators, by contrast, need evidence that can be independently assessed, documented, and linked to product quality, safety, and robustness. Regulatory-facing AI discussions indicate that transparency must be paired with validation and governance, not treated as a purely visual or user-interface function [9, 16].

The resulting translational problem is that AI-generated formulation designs may be algorithmically impressive but institutionally fragile. They can appear innovative to developers while remaining difficult to justify to reviewers, quality teams, or regulators. Table 1 contrasts the current capabilities of generative AI for formulation design with the explainability expectations of scientists and regulators. This gap analysis shows why formulation AI requires a dual-purpose explanation model rather than a generic interpretability checklist [5, 17].

Table 1. AI-Generated Formulation Design versus Explainability Expectations: A Gap Analysis

Dimension

Current AI-generated formulation capability

Scientific explainability expectation

Regulatory explainability expectation

Unresolved gap

Design generation

Models can propose candidate compositions, excipient combinations, and processing conditions from prior data and optimisation objectives.

Scientists need to understand why a proposed design is mechanistically plausible and how it relates to known formulation principles.

Regulators need a documented rationale showing how the design supports quality, safety, and intended performance.

Generated outputs are often stronger as predictions than as explainable formulation arguments.

Search across formulation space

AI can explore high-dimensional spaces more rapidly than conventional trial-and-error screening.

Scientists need interpretable maps of the search space, including trade-offs, constraints, and excluded regions.

Regulators need evidence that the explored design space is justified, bounded, and reproducible.

Search efficiency does not automatically produce auditable design-space knowledge.

Model performance

Deep learning and optimisation models can improve prediction of formulation outcomes when trained on relevant datasets.

Scientists need performance explanations that distinguish meaningful mechanisms from statistical correlations.

Regulators need validation evidence, uncertainty estimates, and robustness testing across relevant use conditions.

Predictive accuracy alone is insufficient for scientific understanding or regulatory confidence.

Local formulation recommendation

AI can recommend a specific formulation candidate for further development.

Scientists need local explanations that identify decisive variables and plausible drug–excipient–process interactions.

Regulators need traceable evidence explaining why that candidate is acceptable within a quality risk framework.

Local recommendations often lack a structured explanation package.

Lifecycle use

AI systems can potentially support iterative optimisation, scale-up, and post-change assessment.

Scientists need explanations that remain meaningful as formulation knowledge evolves.

Regulators need version control, change rationale, and continuity of evidence across lifecycle stages.

Most AI explanations are not designed for lifecycle documentation.

Figure 1 illustrates the central conceptual gap between AI-generated formulation capability and the dual explainability expectations of formulation scientists and regulatory assessors.

Figure 1. Conceptual Gap between AI-Generated Formulation Design Capability and Dual Explainability Expectations

Figure 1. Conceptual Gap between AI-Generated Formulation Design Capability and Dual Explainability Expectations

AI-Generated Formulation Logic

AI-generated formulation design begins with the conversion of pharmaceutical knowledge into machine-readable representations. These representations may include drug descriptors, excipient properties, formulation composition, process parameters, stability outcomes, dissolution behaviour, rheological measures, printability outcomes, or biological performance measures [5, 10, 12]. In chemical design, continuous latent representations enabled models to navigate molecular space and generate new candidates [1]. In formulation science, an analogous challenge is to create representation systems that encode formulation variables without losing the physical meaning of composition, process, and performance relationships.

Generative AI models operate by learning patterns in existing design spaces and then proposing candidates that satisfy selected objectives. Variational autoencoders, reinforcement-learning models, and other generative systems have shown how algorithmic search can move beyond direct imitation of known examples [1, 3, 18]. Although many foundational examples come from molecular design, the same principle can be transferred to formulation design when candidate outputs include quantitative excipient ratios, process settings, and predicted quality attributes. The formulation output is therefore not merely a label or prediction but a proposed product configuration.

Bayesian optimisation and machine learning-guided experimentation add a further logic of iterative decision-making. Instead of generating candidates in a single step, these methods use model predictions and uncertainty estimates to decide which formulation experiments or design regions should be explored next [5, 14]. Robotic and semi-self-driven formulation systems extend this logic by linking prediction, experiment, and feedback into partially automated development loops [19]. Such systems promise acceleration, but they also amplify the need for explanations because each AI-selected experiment can influence the trajectory of product development.

The black-box problem is inherent because formulation design spaces contain nonlinear, conditional, and context-dependent relationships. Deep learning systems in drug discovery and formulation can learn patterns that are difficult to reduce to simple rules, while high-dimensional excipient–drug–process interactions may produce effects that are not obvious from individual variables [4, 20, 21]. This opacity becomes problematic when AI outputs are used to justify development decisions, because a formulation proposal must eventually be translated into experimentally testable reasoning and regulatory documentation. An AI-generated formulation that cannot explain its own plausibility remains incomplete as a scientific and quality claim.

Explainability Requirements

The first explainability requirement is global understanding of how the model behaves across formulation space. Global explanations should identify which formulation variables, physicochemical descriptors, process inputs, and quality attributes most strongly shape model behaviour [6, 7]. For formulation scientists, this helps determine whether the model has learned plausible domain relationships rather than accidental correlations. For regulators, global explanations help clarify whether the system’s operating logic is stable, bounded, and consistent with the intended design context.

The second requirement is local explanation of individual formulation recommendations. Local XAI is especially important because scientific and regulatory decisions are rarely made about a model in the abstract; they are made about a specific formulation candidate, process condition, or change proposal. Explainable AI for drug discovery has already emphasised that the usefulness of an explanation depends on whether it can clarify the rationale behind a concrete prediction [17]. In formulation design, this means explaining why a particular excipient concentration, polymer grade, lipid composition, or processing window was selected for a specific product objective.

The third requirement is mechanistic translation. Formulation scientists need explanations that can be converted into hypotheses about solubility enhancement, release control, stabilisation, manufacturability, or patient-facing performance [5, 12]. In biologics and complex formulations, machine learning may reveal patterns related to developability or formulation robustness, but these patterns must still be interpreted through pharmaceutical science [11]. A mechanistic explanation does not require the model to become a full physical simulator, but it must allow experts to connect model outputs to plausible formulation reasoning.

The fourth requirement is regulatory evidence alignment. AI-generated formulation explanations must support documentation, validation, uncertainty communication, and traceability across development and lifecycle decisions [9, 16]. In high-stakes settings, interpretable models have been argued to be preferable to opaque models when decision consequences are serious, but where complex models are used, the explanation burden becomes greater [22]. Regulatory-grade explainability therefore requires more than visually appealing feature plots. It requires a structured explanation package that can be reviewed, challenged, reproduced, and connected to product quality risk.

The fifth requirement is audience-specific usability without fragmentation. Scientists and regulators need different forms of explanation, but they should not receive disconnected narratives that create contradictory interpretations of the same AI-generated design. Table 2 defines the distinct explainability requirements for scientific and regulatory decision-making. The table frames explainability as a shared infrastructure that must support discovery reasoning, development control, and regulatory assessment simultaneously [15, 23].

Table 2. Explainability Requirements for AI-Generated Formulations: Scientific Understanding Versus Regulatory Evidence

Explainability requirement

Scientific decision-making need

Regulatory decision-making need

Suitable explanation form

Risk if absent

Global model behaviour

Understand the main variables and relationships that shape the model’s formulation logic.

Assess whether the model operates within a defined and justified design context.

Feature importance, sensitivity analysis, design-space visualisation, and model behaviour summaries.

The model may appear accurate but remain scientifically and institutionally opaque.

Local recommendation rationale

Understand why a specific formulation candidate was proposed.

Review the justification for a specific product or process design decision.

Local attribution, counterfactual comparison, nearest-neighbour evidence, and formulation-specific rationale.

Individual AI-generated candidates may be difficult to trust or defend.

Mechanistic plausibility

Convert AI outputs into testable hypotheses about formulation performance.

Determine whether the proposed rationale is consistent with established product and process understanding.

Mechanism-linked explanation narratives, domain-rule overlays, and expert interpretation panels.

AI outputs may become unstructured suggestions rather than knowledge-building decisions.

Uncertainty and robustness

Identify where the model is confident, uncertain, or extrapolating beyond familiar formulation space.

Evaluate validation adequacy, risk boundaries, and the need for additional evidence.

Uncertainty estimates, applicability-domain analysis, stress testing, and robustness summaries.

Overconfident recommendations may lead to unsafe or poorly justified decisions.

Auditability and traceability

Reconstruct how model outputs were generated and how experts used them.

Independently verify data provenance, model version, decision history, and change rationale.

Model cards, data lineage records, version-controlled explanation reports, and decision logs.

AI-supported decisions may fail review because their development history is not reproducible.

Human usability

Enable scientists to query, challenge, and refine model explanations interactively.

Enable assessors to inspect the evidence without relying on developer assertions.

Interactive dashboards, structured reports, and layered explanation summaries.

Users may either reject useful models or over-trust outputs they do not understand.

Scientific Decision-Making

Scientific decision-making with AI-generated formulation designs begins with expert interrogation rather than passive acceptance. Formulation scientists need to know whether an AI proposal is compatible with known drug properties, excipient behaviour, processing constraints, and intended product performance. The broader movement toward automated drug discovery has shown that computational systems can accelerate candidate generation, but scientific value depends on whether generated outputs can be interpreted as knowledge rather than only as ranked suggestions [24]. In formulation development, this means that the scientist must remain an active evaluator of the design rationale.

Interactive explanation is especially important because formulation decisions are rarely based on a single variable. A scientist may need to ask how polymer concentration, surfactant type, lipid ratio, processing temperature, or drying conditions jointly influenced a recommendation. Machine learning-directed formulation development has demonstrated the value of using algorithmic models to navigate complex development spaces, but such models become more useful when their predictions can be linked to formulation trade-offs [5, 25]. Without interaction, explanation risks becoming a static report rather than a tool for scientific reasoning.

Visualisation also has a central role in scientific decision-making because it can make high-dimensional formulation spaces more navigable. Reviews of molecular and pharmaceutical AI show that latent spaces, property maps, and generative trajectories can reveal relationships that are difficult to detect through tabular data alone [26, 27]. For formulation scientists, the relevant visual question is not only where the recommended formulation sits in the design space. It is also whether the formulation lies near known successful regions, near uncertain boundaries, or in a region where the model may be extrapolating.

A major risk is automation bias, in which scientists over-trust an AI-generated proposal because it appears computationally sophisticated. Human-centred explainability research emphasises that explanations should help users calibrate trust rather than simply increase trust [15]. This is crucial in formulation design because a plausible-looking AI output may reflect data bias, incomplete experimental coverage, or a statistical association that lacks formulation meaning. Scientific decision-making therefore requires explanation systems that encourage challenge, comparison, and expert disagreement, not only acceptance.

Regulatory Decision-Making

Regulatory decision-making differs from scientific exploration because it requires independently reviewable evidence. A regulatory assessor cannot rely on the developer’s claim that an AI-generated formulation is optimal; the assessor needs to understand the data, model, assumptions, validation process, uncertainty, and decision rationale. Regulatory-facing AI discussions in clinical and pharmaceutical contexts show that transparency must be linked to accountability and independent assessment [8, 9]. For AI-generated formulations, explanation must therefore become part of the evidence package rather than a supplementary technical appendix.

The regulatory question is not whether AI was used, but whether the use of AI supports or weakens product and process understanding. Machine learning in pharmaceutical development can assist formulation selection, biologics formulation, and quantitative modelling, yet these applications must be connected to a control strategy that defines critical variables and justifies acceptable ranges [11, 12, 16]. An AI-generated formulation must therefore be explained in terms of how it supports quality attributes, manufacturing consistency, and risk control. A model output that cannot be translated into such terms remains difficult to evaluate in a regulatory context.

Regulatory-grade explainability should also address reproducibility and lifecycle management. AI-generated designs may evolve as new formulation data, process experience, or post-approval information become available, which means that explanation systems must preserve model version, data provenance, and decision history. The growing use of artificial intelligence across drug development, manufacturing, quality control, and post-market surveillance has increased the importance of documentation that follows a product across its lifecycle [28]. In this context, explanation is not a one-time interpretation but a continuing record of why decisions were made.

AI-generated formulation designs must also be aligned with quality-by-design reasoning, even when formal regulatory guidance for AI explanation remains incomplete. The logic of QbD requires product and process understanding, risk-based control, and justification of design space; AI explanations should therefore clarify how model recommendations relate to formulation variables, critical quality attributes, and control strategy. Quantitative modelling perspectives in drug development emphasise that best practices require transparent assumptions, fit-for-purpose validation, and communication of model limitations [16]. These principles provide a foundation for explaining AI-generated formulation proposals in a form that regulators can assess.

Transparency Barriers

The first transparency barrier is technical opacity. Deep learning and generative models can learn nonlinear patterns across heterogeneous data, but their internal representations may not correspond directly to familiar formulation concepts [4, 20]. In molecular design, latent representations and generative search can produce useful candidates without making the design logic immediately interpretable [1, 2]. In formulation design, this opacity is amplified because the relevant performance outcome may depend on coupled drug–excipient–process relationships rather than on a single molecular structure.

The second barrier is the high dimensionality and incompleteness of formulation data. Formulation datasets are often smaller, less standardised, and more context-dependent than the datasets used in many image or language applications. Machine learning studies in drug formulation and biologics development show that useful models can be built, but their reliability depends heavily on data quality, domain coverage, and careful interpretation [5, 11, 13]. When the underlying dataset does not adequately represent the formulation design space, an explanation may describe the model’s learned correlations without revealing the limits of those correlations.

The third barrier is methodological mismatch between generic XAI tools and pharmaceutical formulation needs. SHAP-type attribution, surrogate models, counterfactuals, and attention-based explanations may identify influential inputs, but they do not automatically explain formulation mechanisms or regulatory relevance [6, 7, 17]. For example, an attribution method may show that a surfactant concentration influenced a predicted stability outcome, but it may not explain whether the effect reflects interfacial stabilisation, micellar solubilisation, viscosity change, or confounding in the training data. This creates a gap between computational interpretability and pharmaceutical intelligibility.

The fourth barrier is organisational and regulatory uncertainty. Developers may lack expertise in XAI, quality teams may lack procedures for reviewing AI-generated formulation evidence, and regulators may lack agreed criteria for what constitutes an acceptable explanation. High-stakes AI literature warns that opaque models can be inappropriate when decision consequences are serious, while regulatory-facing AI work shows that institutional trust depends on auditability and validation [9, 22]. Without shared expectations, AI-generated formulations may be over-explained in technically irrelevant ways or under-explained in ways that fail regulatory review.

Proposed Explainability Framework

The proposed framework treats explainability as a layered decision architecture rather than a single method. The first layer is global model explanation, which describes how the AI system behaves across formulation space and which variables most strongly shape its recommendations. This layer should include sensitivity analysis, feature attribution, design-space mapping, and applicability-domain assessment, drawing on established XAI principles while adapting them to formulation variables [6, 7]. Its function is to show whether the model has learned a plausible and bounded formulation logic.

The second layer is local formulation explanation, which explains why a particular AI-generated formulation was proposed. Local explanations should identify the decisive inputs, compare the candidate with nearby alternatives, describe counterfactual changes that would alter the prediction, and indicate whether the candidate lies within familiar or uncertain design regions. Explainable AI for drug discovery has shown that local interpretability is essential when users must evaluate specific model outputs rather than general model behaviour [17]. In formulation design, this local explanation should become the core bridge between AI recommendation and expert scientific review.

The third layer is mechanistic and scientific translation. This layer converts computational explanations into formulation hypotheses that scientists can test and refine. It should connect AI-selected variables to possible mechanisms such as solubility enhancement, release control, physical stabilisation, manufacturability, or patient-facing performance, while clearly distinguishing confirmed knowledge from model-derived hypotheses. Machine learning formulation studies demonstrate that algorithmic systems can guide development, but their scientific usefulness increases when outputs can be interpreted through pharmaceutical principles [12, 21, 25].

The fourth layer is regulatory evidence packaging, which translates explanation into documentation, traceability, validation, uncertainty communication, and lifecycle decision records. This layer should include model development records, training-data provenance, version history, validation protocols, robustness testing, expert review logs, and formulation-specific rationale. Table 3 presents the proposed framework for dual-purpose explainability in pharmaceutical formulation AI. The framework integrates global interpretability, local rationale, mechanistic translation, and regulatory evidence so that one AI-generated formulation design can support both scientific understanding and regulatory assessment [16, 28].

Table 3. Proposed Explainability Framework for AI-Generated Formulation Designs: Components, Tools, and Implementation

Framework component

Primary purpose

Suitable tools or methods

Scientific decision value

Regulatory decision value

Global model explanation layer

Explain how the model behaves across the formulation design space.

Feature attribution, sensitivity analysis, surrogate models, applicability-domain mapping, and design-space visualisation.

Helps scientists identify learned formulation patterns, influential variables, and possible trade-offs.

Helps assessors understand model scope, operating boundaries, and general decision logic.

Local formulation explanation layer

Explain why a specific formulation candidate was generated or recommended.

Local attribution, counterfactual comparison, nearest-neighbour comparison, uncertainty estimate, and candidate-specific explanation report.

Helps scientists judge whether a proposed formulation is plausible and worth experimental follow-up.

Helps assessors review the rationale for a specific product or process decision.

Mechanistic translation layer

Convert computational signals into pharmaceutical formulation hypotheses.

Domain-rule overlays, expert annotation, mechanism mapping, and hypothesis-generation templates.

Supports scientific learning by linking AI outputs to drug–excipient–process mechanisms.

Supports product and process understanding by connecting AI rationale to quality attributes.

Uncertainty and robustness layer

Identify confidence, extrapolation, and vulnerability of model recommendations.

Prediction intervals, ensemble disagreement, stress testing, out-of-domain detection, and robustness summaries.

Helps scientists avoid over-trusting weak or extrapolated recommendations.

Helps assessors evaluate risk, evidence sufficiency, and need for confirmatory data.

Auditability and governance layer

Preserve the evidence trail behind AI-supported formulation decisions.

Data lineage, model cards, version control, decision logs, validation reports, and expert-review records.

Helps development teams reconstruct and improve decision pathways.

Enables independent review, traceability, and lifecycle change management.

Human-interaction layer

Allow users to query, challenge, and refine explanations.

Interactive dashboards, explanation queries, scenario testing, and structured expert feedback.

Promotes calibrated trust and domain-informed interpretation.

Demonstrates that human oversight is embedded in AI-supported decision-making.

Figure 2 presents the proposed dual-purpose explainability framework that converts AI-generated formulation outputs into scientific understanding and regulatory evidence.

Figure 2. Dual-Purpose Explainability Framework for AI-Generated Pharmaceutical Formulation Designs

Figure 2. Dual-Purpose Explainability Framework for AI-Generated Pharmaceutical Formulation Designs

Implementation Pathway

The first implementation step is to develop XAI tools that are specific to formulation data rather than borrowed unchanged from generic machine-learning workflows. These tools should represent excipient function, composition ranges, process variables, critical quality attributes, and performance outcomes in ways that formulation scientists can interpret. Existing work on AI-guided formulation, 3D printing, suspensions, vaccines, and robotic formulation shows that the technical basis for AI-supported development is expanding rapidly [10, 13, 14, 19]. The next step is to ensure that each AI recommendation is accompanied by an explanation that is scientifically meaningful and not only computationally valid.

The second step is to conduct pilot studies with formulation scientists, quality specialists, and regulatory assessors. These studies should examine whether explanation packages actually improve decision quality, reduce inappropriate trust, and make AI-generated designs easier to audit. Human-centred XAI research shows that explanation quality must be evaluated in relation to the user’s task, expertise, and decision environment [15, 23]. For pharmaceutical formulation, this means testing explanations in realistic development scenarios where experts must accept, reject, modify, or justify an AI-generated formulation proposal.

The third step is to integrate explainability expectations into emerging governance structures for AI in pharmaceutical development. Industry and regulatory stakeholders should define minimum documentation standards for AI-generated formulation designs, including data provenance, model validation, uncertainty communication, local rationale, and lifecycle traceability. The increasing use of artificial intelligence across pharmaceutical development makes this integration urgent because AI may soon influence formulation design, manufacturing strategy, quality control, and post-market change management in connected ways [16, 28]. Explainability should therefore be treated as part of responsible pharmaceutical innovation rather than as a late-stage regulatory defence.

Figure 3 outlines the implementation pathway for embedding dual-purpose explainability into AI-driven formulation development and regulatory preparation.

Figure 3. Implementation Pathway for Regulatory-Ready Explainable AI in Pharmaceutical Formulation Development

Figure 3. Implementation Pathway for Regulatory-Ready Explainable AI in Pharmaceutical Formulation Development

Conclusion

AI-generated formulation design has the potential to transform pharmaceutical development by expanding design-space exploration, accelerating optimisation, and revealing formulation possibilities that conventional workflows may overlook. Yet this potential will remain limited if generated designs cannot be explained in ways that scientists and regulators can use. The central challenge is not only to make AI more powerful, but to make its outputs interpretable, auditable, and accountable.

This article proposed a dual-purpose explainability framework for AI-generated pharmaceutical formulation designs. The framework distinguishes scientific understanding from regulatory evidence while showing that both must be supported by the same transparent decision infrastructure. Global model explanation, local formulation rationale, mechanistic translation, uncertainty communication, and regulatory evidence packaging together provide a pathway from black-box formulation prediction to explainable formulation intelligence.

The future of AI-driven formulation science will depend on collaboration among formulation scientists, AI developers, quality experts, regulators, and industry decision-makers. Explainability must be embedded early in model development, experimental design, and regulatory planning rather than added after the AI system has already produced recommendations. Only then can AI-generated formulation designs become not merely computationally innovative, but scientifically credible and regulatory-ready.

Acknowledgements

None

Conflict of interest

None

Financial support

None

Ethics statement

None

References

Gómez-Bombarelli R, Wei JN, Duvenaud D, Hernández-Lobato JM, Sánchez-Lengeling B, Sheberla D, et al. Automatic chemical design using a data-driven continuous representation of molecules. ACS Cent Sci. 2018;4(2):268-76.
Sanchez-Lengeling B, Aspuru-Guzik A. Inverse molecular design using machine learning: generative models for matter engineering. Science. 2018;361(6400):360-5.
Popova M, Isayev O, Tropsha A. Deep reinforcement learning for de novo drug design. Sci Adv. 2018;4(7):eaap7885.
Chen H, Engkvist O, Wang Y, Olivecrona M, Blaschke T. The rise of deep learning in drug discovery. Drug Discov Today. 2018;23(6):1241-50.
Bannigan P, Aldeghi M, Bao Z, Häse F, Aspuru-Guzik A, Allen C. Machine learning directed drug formulation development. Adv Drug Deliv Rev. 2021;175:113806.
Guidotti R, Monreale A, Ruggieri S, Turini F, Giannotti F, Pedreschi D. A survey of methods for explaining black box models. ACM Comput Surv. 2018;51(5):1-42.
Arrieta AB, Díaz-Rodríguez N, Del Ser J, Bennetot A, Tabik S, Barbado A, et al. Explainable Artificial Intelligence (XAI): concepts, taxonomies, opportunities and challenges toward responsible AI. Inf Fusion. 2020;58:82-115.
Benjamens S, Dhunnoo P, Meskó B. The state of artificial intelligence-based FDA-approved medical devices and algorithms: an online database. NPJ Digit Med. 2020;3(1):118.
Muehlematter UJ, Daniore P, Vokinger KN. Approval of artificial intelligence and machine learning-based medical devices in the USA and Europe (2015–20): a comparative analysis. Lancet Digit Health. 2021;3(3):e195-203.
Elbadawi M, Castro BM, Gavins FK, Ong JJ, Gaisford S, Pérez G, et al. M3DISEEN: a novel machine learning approach for predicting the 3D printability of medicines. Int J Pharm. 2020;590:119837.
Narayanan H, Dingfelder F, Butté A, Lorenzen N, Sokolov M, Arosio P. Machine learning for biologics: opportunities for protein engineering, developability, and formulation. Trends Pharmacol Sci. 2021;42(3):151-65.
Narayanan H, Dingfelder F, Condado Morales I, Patel B, Heding KE, Bjelke JR, et al. Design of biopharmaceutical formulations accelerated by machine learning. Mol Pharm. 2021;18(10):3843-53.
Zulbeari N, Wang F, Mustafova SS, Parhizkar M, Holm R. Machine learning strengthened formulation design of pharmaceutical suspensions. Int J Pharm. 2025;668:124967.
Li L, Back SI, Ma J, Guo Y, Galeandro-Diamant T, Clénet D. Bayesian optimization and machine learning for vaccine formulation development. PLoS One. 2025;20(6):e0324205.
Markus AF, Kors JA, Rijnbeek PR. The role of explainability in creating trustworthy artificial intelligence for health care: a comprehensive survey of the terminology, design choices, and evaluation strategies. J Biomed Inform. 2021;113:103655.
Terranova N, Renard D, Shahin MH, Menon S, Cao Y, Hop CE, et al. Artificial intelligence for quantitative modeling in drug discovery and development: an innovation and quality consortium perspective on use cases and best practices. Clin Pharmacol Ther. 2024;115(4):658-72.
Jiménez-Luna J, Grisoni F, Schneider G. Drug discovery with explainable artificial intelligence. Nat Mach Intell. 2020;2(10):573-84.
Polykovskiy D, Zhebrak A, Sanchez-Lengeling B, Golovanov S, Tatanov O, Belyaev S, et al. Molecular sets (MOSES): a benchmarking platform for molecular generation models. Front Pharmacol. 2020;11:565644.
Ros H, Abdalla Y, Cook MT, Shorthouse D. Efficient discovery of new medicine formulations using a semi-self-driven robotic formulator. Digital Discovery. 2025;4(8):2263-72.
Lavecchia A. Deep learning in drug discovery: opportunities, challenges and future prospects. Drug Discov Today. 2019;24(10):2017-32.
Bao Z, Bufton J, Hickman RJ, Aspuru-Guzik A, Bannigan P, Allen C. Revolutionizing drug formulation development: the increasing impact of machine learning. Adv Drug Deliv Rev. 2023;202:115108.
Rudin C. Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead. Nat Mach Intell. 2019;1(5):206-15.
Gunning D, Aha DW. DARPA's explainable artificial intelligence program. AI Mag. 2019;40(2):44-58.
Schneider G. Automating drug discovery. Nat Rev Drug Discov. 2018;17(2):97-113.
Murray JD, Lange JJ, Bennett-Lenane H, Holm R, Kuentz M, O'Dwyer PJ, et al. Advancing algorithmic drug product development: recommendations for machine learning approaches in drug formulation. Eur J Pharm Sci. 2023;191:106562.
Elton DC, Boukouvalas Z, Fuge MD, Chung PW. Deep learning for molecular design—a review of the state of the art. Mol Syst Des Eng. 2019;4(4):828-49.
Vamathevan J, Clark D, Czodrowski P, Dunham I, Ferran E, Lee G, et al. Applications of machine learning in drug discovery and development. Nat Rev Drug Discov. 2019;18(6):463-77.
Huanbutta K, Burapapadh K, Kraisit P, Sriamornsak P, Ganokratanaa T, Suwanpitak K, et al. Artificial intelligence-driven pharmaceutical industry: a paradigm shift in drug discovery, formulation development, manufacturing, quality control, and post-market surveillance. Eur J Pharm Sci. 2024;203:106938.

Author information

George Papadopoulos & Eleni Georgiou contributed to this work.

Authors and affiliations

Department of Pharmaceutical Innovation and Systems, Faculty of Pharmacy, National and Kapodistrian University of Athens, Athens, Greece
George Papadopoulos & Eleni Georgiou

Corresponding author

Correspondence to George Papadopoulos

Rights and permissions

Open Access The author(s) retain copyright. This article is licensed under the Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License. It may be shared and adapted for non-commercial purposes with appropriate attribution, an indication of changes, and distribution of adaptations under the same license. Third-party material may be subject to separate terms identified in its credit line. View the license at https://creativecommons.org/licenses/by-nc-sa/4.0/.

About this article

Cite this article

Vancouver
Papadopoulos G, Georgiou E. Explainable AI-Generated Formulation Designs for Regulatory and Scientific Decision-Making. . 0;0:192.
APA
Papadopoulos, G., & Georgiou, E. (0). Explainable AI-Generated Formulation Designs for Regulatory and Scientific Decision-Making. EAMD 3, 0, 192.
Received
14 October 2025
Revised
25 November 2025
Accepted
15 December 2025
Published
10 January 2026
Version of record
10 January 2026

Share this article

Easily share this article with others using the link below:

Explainable AI-Generated Formulation Designs for Regulatory and Scientific Decision-Making
Scan to access
this article

Ready to submit?
Start a new submission or continue a submission in progress:
Submission Portal Author Guidelines

Follow this journal
Get notified of new updates and articles.