Conceived and designed the experiments: FC GG MAB JNC NHP MMH. Performed the experiments: FC CBM AB MMH. Analyzed the data: FC GG MMH. Contributed reagents/materials/analysis tools: MMH. Wrote the paper: FC GG MAB AB JNC NHP MMH.
We consider the problem of optimizing a liposomal drug formulation: a complex chemical system with many components (e.g., elements of a lipid library) that interact nonlinearly and synergistically in ways that cannot be predicted from first principles.
The optimization criterion in our experiments was the percent encapsulation of a target drug, Amphotericin B, detected experimentally via spectrophotometric assay. Optimization of such a complex system requires strategies that efficiently discover solutions in extremely large volumes of potential experimental space. We have designed and implemented a new strategy of evolutionary design of experiments (Evo-DoE), that efficiently explores high-dimensional spaces by coupling the power of computer and statistical modeling with experimentally measured responses in an iterative loop.
We demonstrate how iterative looping of modeling and experimentation can quickly produce new discoveries with significantly better experimental response, and how such looping can discover the chemical landscape underlying complex chemical systems.
Formulation of a stable lipid membrane with good drug complexation characteristics is a complex experimental problem as there are many possible components (qualitative variables) with a wide range of relative amounts (quantitative variables) in the recipe. This makes drug formulation a good test case for methodologies for discovering and optimizing complex chemical systems. The experiments we present have five qualitative factors and nine components whose concentrations are varied in the aqueous phase, to form an experimental space that contains a total of 82,950 possible experiments.
The experimental approach to the optimization problem considered here is iterated high-throughput experimentation. Many experiments are performed in parallel (e.g., using a 96-well plate format), the results are analyzed, and a successive generation of experiments is performed, with the goal of progressively better results with each generation. After each generation of experiments, the experimenter is faced with the problem of designing a limited yet informative set of experiments from an expansive experimental space for the successive generation.
We address the problem of designing experiments for the exploration of high-dimensional experimental spaces by applying statistical modeling and predictive methods to the entire experimental space at each generation. We build models from the raw experimental data, and we use those models to choose where next to sample the experimental space. This methodology is iterated as illustrated in
Note the predictive modeling procedure in the loop.
To start the optimization process, we sparsely sample the experimental space with a random selection of experiments. We build models of the desired response from the experimental data, design the next sparse sampling of the experimental space, and the process repeats. Tightly coupling experiments with statistical modeling and predictive algorithms enables successful optimization of desired chemical behavior of supramolecular structures in large high-dimensional experimental spaces.
The problem of designing each successive generation of experiments is difficult for two related reasons: the statistical uncertainty of any predictive model given the relatively small amount of data (∼102–103 samples), and the size of the space (particularly when the number of dimensions (experimental variables) is large). Large regions of the space are inevitably unrepresented in the experimental data. To address the first difficulty, we use a bootstrapping procedure in the model-building process to ensure that the model's complexity is not excessive given the quantity of data available. The second difficulty is handled by combining model-based prediction of experiments with a weighted random sampling of the experimental space, where the distribution for the random sample is biased toward the unsampled regions of the space. Given these two difficulties, our approach to the experimental design problem may be seen as an example of establishing an
Traditional approaches to DoE for high-dimensional experimental spaces use a “screening design” to identify a few significant factors of many potential ones, or to identify a few significant factors that embody the “main effects” that are presumed to be an order of magnitude more important than “interaction effects”
Amphotericin B is an antifungal drug used to treat systemic fungal infections. The drug can cause nephrotoxicity if present in high doses. However, when intercalated within a lipid membrane, the hydrophobic drug can be administered to a patient with minimal toxic effects. There are currently three different lipid formulations of Amphotericin B on the market
In experiments presented here, we describe an Evo-DoE procedure to search within a defined space of 82,950 possible experiments for lipid combinations that maximize the amount of Amphotericin B entrapped in the formulation. In a space of this size and complexity, exhaustive screening is impractical, conventional DoE has severe difficulties handling the multiple qualitative variables, and simple hill-climbing approaches tend to fail because of the interaction effects in the system. Within a few iterations (exploration of ∼0.5% of the space) we found many new promising Amphotericin B formulations not previously reported. In addition we were able to determine experimentally the response surface for the lipid-drug combinations.
Evo-DoE is an iterated, cyclical process involving experimentally measuring responses of sparsely sampled recipes from a space of possible experimental recipes. One approach to implementing evolutionary DoE is to use a genetic algorithm, which assigns each possible experiment a genetic code, and then progresses from one generation of experiments to the next by applying genetic operators analogous to mutation and crossover to winning recipes
The Evo-DoE process used here started with a set of 90 randomly selected recipes that included all the pairwise combinations of lipids in our lipid library. The Nth iteration of the Evo-DoE process consisted of the following steps:
Measure the experimental response of the Nth set of recipes;
Optimize the metaparameters of the Nth model, as described below;
Build a model of the entire response surface from the experimentally measured responses for the first N sets of recipes;
Randomly choose 48 recipes from the untried recipes with the top quartile of responses predicted by the model, and add those recipes to the N+1th set of recipes;
Add 12 more randomly selected untried recipes to the N+1th set of recipes.
Neural network metaparameters are often optimized by a “bootstrapping” process
The response of each recipe was calculated as follows: Three spectrophotometric absorbance measurements were taken from experimental replicates in three separate wells of a 96-well plate at three different times, and the response of a recipe was defined as the average of those nine measurements.
The system was quickly optimized (
Error bars on the best new recipe were taken from three repeats performed on the same day (space repeats, see
More detailed analysis of the high response optimum found by Evo-DoE is shown in
For each PG-type lipid, shown in bold, the horizontal section lists the lipids from the group with a net negative charge and the vertical section lists the reagents in the aqueous phase and their corresponding pH values. Response levels (the UV/Vis absorbance of Amphotericin B associated with the formulation): dark grey, >0.20; medium grey, 0.15–0.20; light grey, 0.10–0.15; white, <0.10; Blank cells, not determined. For abbreviations, see
The abundance of every library component for each successive generation of Evo-DoE was recorded. The results shown in
The representation of each lipid, computed as a percentage of recipes that have each particular component, for each generation of Evo-DoE is shown for seven successive generations. For generation 2–7, only model-based recipes are considered. A) The PG lipid group; B) the negatively charged lipid group. See
We chose a sampling of winning recipes found by Evo-DoE based on high response and diversity of components. We then analyzed this group of twelve recipes for the formation of vesicles and stability over time at varying temperatures, and normalized each result based on the standard recipe. It should be noted that we based our response function here solely upon the amount of Amphotericin B associated with the lipids and not on these secondary criteria. However, a response function could in principle combine all three criteria.
The results of the analyses of the winning recipes are shown in
| Recipe | Lipid phase | Aqueous phase | Response | Internal Volume | Stability at 4°C | Stability at 25°C | Stability at 50°C |
| Std | DSPG, Chol, DOPC | Succinic Acid 10 mM, pH 4.5, Na(OH), Sucrose 9%. | 1.0 | 1.00 | 1.0 | 1.0 | 1.0 |
| 1 | PS, Chol, SM | Bicine 100 mM, pH 8.5, Na(OH), Glucose 9%. | 1.8 | 0.50 | 1.1 | 3.7 | 5.0 |
| 2 | DSPG, Chol, DOPM | Bicine 100 mM, pH 8.5, Na(OH), Sucrose 4.5%. | 1.7 | N.A. | 1.11 | 3.7 | 4.3 |
| 3 | DSPG, Ergo, Lino | Mes 100 mM, pH 7, Na(OH), Sucrose 9%. | 1.3 | 0.98 | 2.0 | 5.3 | 6.7 |
| 4 | DSPG, Chol, DOPM | Bicine 10 mM, pH 8.5, Na(OH), Sucrose 9%. | 1.8 | 0.52 | 2.0 | 4.7 | 5.0 |
| 5 | DSPG, Chol, DOPM | Mes 10 mM, pH 7, Na(OH), Sucrose 4.5%. | 1.7 | N.A. | 1.8 | 5.7 | 6.0 |
| 6 | DSPG, Chol, Oleic | Bicine 100 mM, pH 8.5, Na(OH), Glucose 4.5%. | 2.0 | 0.40 | 2.2 | 6.7 | 6.7 |
| 7 | DMPG, Chol, DOPM | Bicine 100 mM, pH 8.5, Na(OH), Glucose 9%. | 2.2 | N.A. | 2.0 | 3.0 | 4.7 |
| 8 | DMPG, Chol, DOPM | Bicine 100 mM, pH 8.5, Na(OH), Glucose 4.5%. | 1.9 | 0.26 | 1.4 | 3.3 | 5.0 |
| 9 | DPPC, Ergo, SM | Bicine 10 mM, pH 8.5, Na(OH). | 1.7 | N.A. | 2.1 | 6.7 | 5.0 |
| 10 | PC, Chol, PS | Bicine 100 mM, pH 8.5, Na(OH), Glucose 9%. | 2.0 | 0.20 | 2.0 | 5.7 | 5.7 |
| 11 | PC, Chol, Lino | Bicine 100 mM, pH 8.5, Na(OH), Glucose 9%. | 2.0 | N.A. | 2.1 | 5.7 | 6.7 |
| 12 | DOPC, Chol, SM | Bicine 100 mM, pH 8.5, Na(OH), Glucose 4.5%. | 1.9 | N.A. | 1.0 | 3.3 | 4.3 |
The experimental space of this system was large and complex, consisting of a 5-factor qualitative/mixture design in lipids crossed with a 7×2 factor design in the aqueous phase. There were 82,950 possible combinations, far exceeding the throughput of the experimental system. Traditional DoE methods handle such complex systems with great difficulty. They were designed to work with small numbers of experimental points when interactions among the system components are weak enough that the space can be simplified
Genetic algorithms have been used in iterated high throughput optimization
It is important to note that the predictive model of the Evo-DoE method presented here is a statistical model built entirely from the available experimental data, and contains no expert prior knowledge regarding reactions or self-assembly processes. Neither does it contain structural information as in QSAR models
This refined method of model-based screening resulted in rapid and effective location of an optimum formulation of vesicle-encapsulated Amphotericin B. The optimum was found using only 0.5% of the total possible points, and so may be only a local optimum, but coverage of the space was nevertheless quite broad. Many formulation variations are possible and available for characterization and development into commercially viable products.
It is also possible to get a rough picture of the response landscape dictated by the lipid and aqueous phase combinations and their interactions with the target. The high dimensionality and qualitative nature of the factors makes conventional visualization difficult. We found it best to section the measured response surface to get intuitions about its topology, as shown in
The solutions found in our system may also be influenced by the particular high throughput method we employed, as new recipes were selected from previous ones based on how well the latter complexed with Amphotericin B using our protocol. Other protocols may place other selective factors or stresses on the system and produce different optima. However, we found many different types of recipes that can be used for Amphotericin B formulation. Other groups using different protocols have also found some of the factors that contributed to a positive response in our system. For example, the specific favorable interaction between Amphotericin B and the lipid DSPG has been known for some time, and forms the basis for the Ambisome patent (
Although our response function did not include other important parameters for drug formulation development, such as structure of the resulting particles, stability of the formulation over time, or pharmacokinetics, these could be included in a response function. The predictive Evo-DoE framework can accommodate the incorporation of such parameters. We have characterized the structure and stability of the fittest drug formulations post-optimization, and see variation in these parameters. Using our protocol with a multi-component response function would allow direct optimization of drug formulations, from the initial incorporation of drug into lipid structures through to animal models.
Our Evo-DoE methodology “closes the loop” in iterated adaptive experimentation in a high-dimensional experimental space. This closed loop can be fully automated if autonomously operating robots are used to conduct high-throughput experimentation, with the results fed directly into computers for automated statistical analysis of experimental results, and then automated intelligent design of the subsequent round of experiments could be fed directly back into the experimental robot platform. This can ultimately lead to 24/7 operation and rapid optimization of complex experimental spaces, and would mark a significant milestone in automating the scientific process
The task of discovering and optimizing complex chemical systems suffers from two recurring problems: First, beneficial nonlinear interactions among system components cannot be inferred from basic chemical and physical laws, and second, as the number of system constituents and experimental parameters increases, traditional screening of experimental spaces becomes impractical and economically unfeasible. The general method applied here to optimize liposomal formulations of Amphotericin B, predictive Evo-DoE, can be used generally to discover and optimize other kinds of complex chemical systems, thus yielding a new tool for solving the problem of chemical complexity.
Water, mes, phosphate buffered saline (PBS), glutamic acid (Glu), succinic acid (Succ), hepes, bicine and trizma (Tris) were purchased from Sigma Aldrich. Sodium hydroxide was purchased from Merck.
Chloroform, ethanol, methanol and dimethyl sulfoxide (DMSO) were purchased from Sigma Aldrich.
Egg PC (L-α-phosphatidylcholine, hydrogenated), POPC (1-palmitoyl-2-oleoyl-
Amphotericin B from Streptomyces ssp. in powder form and β-octylglucopyranoside (BOG) in powder form was purchased from Fluka. Sepharose 4B, Rhodamine 6G, calcein, glucose and sucrose were purchased from Sigma.
Buffers of stock solution, 100 mM, pH 4.5–7 and 8.5, NaOH were prepared and stored at 4°C. Buffer titrations were performed with a solution of 10M NaOH in water to adjust pH.
A library of 75 buffers was required to cover all possible combinations of 7 different acids at 2 concentration levels, 3 pHs, and 2 sugars at 2 concentration levels. All were stored at 4°C and used at room temperature (23°C).
The BOG detergent powder was dissolved in water to a 20% (w/v) final concentration. The solution was stored at 4°C.
Stock solutions of phospholipids in HCCl3, HCCl3/MeOH 3∶1 or HCCl3/MeOH 95∶5 were prepared and stored under nitrogen in Chromacol screw cap vials with silicon/teflon septa (Microcolumn Srl), and stored at −20°C. DOPC, POPC, DPPC, PC (egg), SM, PS, DOPM, oleic acid, myristoleic acid, linoleic acid, cholesterol and ergosterol were prepared at 5 mM and diluted when needed to 1 mM. CL, DPPG, POPG and DMPG were prepared at 6 mM and diluted to 0.6 mM. DSPG in powder was prepared at 0.6 mM.
Amphotericin B powder was dissolved in DMSO at 3.3 mM. A clear orange solution was obtained and stored at 4°C.
Absorbance measurements at 415nm in the high-throughput experiment were recorded with a PerkinElmer Wallace 1420 Victor 3 Multilabel Counter, designed for 96-well plates. Absorbance spectra of the Amphotericin B in aqueous solution and in solution with vesicles were recorded with a PerkinElmer Lambda 25 spectrophotometer, using a quartz cuvette.
The high-throughput experiment was performed with a robotic workstation for liquid handling, Xiril 75-1-2 (Switzerland). The hardware layout was designed specifically for the experimental protocol developed. The library of buffers, DMSO and BOG were contained in Chromacol 10 SV tubes and placed in a removable stainless steel tube decktray for the Xiril 75 with 96 12×75 mm positions. The dilution steps of the protocol were performed in 1.5 mL Eppendorf tubes, set in 32 position racks for 1.5 or 2 mL microfuge tubes with lids. At the end of the experiments the racks containing the buffers, DMSO and BOG were stored at 4°C. The custom racks were purchased from Xiril. The robot workstation uses Rainin tips GPS-L250 space saver 960PZ purchased from Elettrofor Sas, Rovigo, Italy.
The following section describes the chemical high-throughput experiment protocol for preparing and analyzing formulations containing Amphotericin B. The DMSO solution containing Amphotericin B (crystallized at 4°C) is warmed on a standard heatblock until it becomes liquid. 2.87 µl of the Amphotericin B solution are transferred by pipette into the bottom of 500 µl volume glass vials (Chromacol 05-MTV-96) and then set in a 96-position deep well vial holder (Chromacol 05-MTP-96) from Microcolumn Srl.
The well plate containing the Amphotericin B vials is then placed in the robotic workstation and the lipids in organic solvent were added to the vials according to a well map, which specifies the exact qualitative and quantitative composition of the resulting mixture in each of the wells. At the end of the distribution step, each vial contains a mixture of three lipids and Amphotericin B.
The resulting plate containing the vials of organic solvent, lipids and Amphotericin B is then transferred into a vacuum chamber connected with a vacuum pump (KNF LAB, Pressure min 1.0 bar) with Teflon membranes. A heatblock at 100°C is set under the well plate for 20 minutes before it is removed and the evaporation process is allowed to continue for another 40 minutes, achieving a thin yellow film on the glass surface of the vials. The well plate is then transferred again to the robotic workstation and 200 µl of the hydration aqueous phase are added according to a well map. The final concentration of the lipids in the mixture is always 500 µM.
Once the lipid film is hydrated it is sonicated at 25°C for 10 minutes (Bandelin Sonorex Digitec), to promote the formation of small unilamellar vesicles (SUV). The robot then transfers 140 µl of the sonicated vesicle solution into Eppendorf tubes, prefilled with 240 µl of buffer. The samples are then centrifuged (Heraeus Biofuge Pico) at 13×103 rpm for 15 minutes to separate the Amphotericin B entrapped in the formulation from free drug crystals in solution. A yellow pellet representing the free drug precipitate is visible on the bottom of the Eppendorf tubes.
178.5 µl of the supernatant is pipetted carefully into new 1.5 ml Eppendorf tubes and repositioned in the robot. To each sample, the robot adds 21 µl of DMSO and 10.5 µl of BOG, reaching a total volume of 210 µl in order to destroy any vesicles in the sample, leaving the Amphotericin B free in the organic solvent to be analyzed spectroscopically.
The robot transfers 200 µl of the resulting solution into a transparent 96-position micro-titration plate. The well plate is covered and stored at 4°C overnight. Before measuring the absorbance of the solution with the spectrophotometer, any bubbles in the samples are removed with a gentle stream of nitrogen. Triplicate absorbance measurements are made at 415 nm.
Formulations containing Amphotericin B based on twelve selected top recipes and the standard were prepared as above. The formulations were then split and incubated in parallel at 4°C, 25°C, and 50°C, over the course of 30 days. For analysis, an aliquot of 10 µl was placed on a slide with 10 µM Rhodamine 6G and analyzed with fluorescent microscopy (Nikon TE2000-S inverted fluorescence microscope with Photometrics Cascade II 512 camera and in-house software). The samples were visualized with a 40× objective and 10 different frames were captured per sample. The captured frames were then analyzed for the presence of aggregates as evidenced by small dye-containing lipid particles. They were scored as according to the criteria: score of 0 for no aggregates seen, 1 for at least one aggregate in only one captured frame, 2 for one aggregate in more than one independent frame, 3 for at least one aggregate in all frames, 4 for many aggregates in all frames. These scores were then averaged over the time course of the experiment and normalized to the standard.
Formulations containing Amphotericin B based on selected top recipes and the standard were prepared as above, but with 10uM calcein in the hydration buffer. The unencapsulated external dye was then separated from the lipid formulations using size exclusion chromatography (sepharose 4B column, 6×1cm). The percent dye encapsulated was then quantified on the PerkinElmer Wallace 1420 Victor 3 Multilabel Counter. The values were then normalized to the standard.
Response values are calculated from the absorbance of Amphotericin B after the removal of the uncomplexed Amphotericin B crystals in solution and after the destruction of the retained vesicle formulation with BOG as detergent. The Amphotericin B absorbance spectrum shows several peaks depending on the physical state of the Amphotericin B
The full experimental space of this experiment contained 82,950 possible recipes. It consisted of all possible combinations of the following two libraries:
An aqueous phase library (
A lipid library (
| Buffers | pH group | Levels (mM) | Sugars | Levels (%) |
| HEPES | 7 | 10, 100 | Glucose | 4.5, 9 |
| MES | 7 | 10, 100 | Sucrose | 4.5, 9 |
| Glutamic acid | 8.5 | 10, 100 | No Sugar | |
| Succinic acid | 4.5 | 10, 100 | ||
| Bicine | 8.5 | 10, 100 | ||
| Tris | 8.5 | 10, 100 | ||
| PBS | 7 | 10, 100 | ||
| No buffer |
| Lipid Categories | Lipids |
| PC | DOPC, POPC, DPPC, PC (egg) |
| PG | CL, DSPG, DPPG, POPG, DMPG |
| Sterol | Cholesterol, Ergosterol |
| Negatively charged | PS, DOPM, oleic acid, myristoleic acid, linoleic acid |
| Sphingo | Sphingomyelin |
Because the lipid combinations form a mixture system, the range of quantitative possibilities (relative concentrations of each of the lipid types) is a polytope. The polytope and its extreme vertices were generated using JMP software, and a set of the four vertices, two midpoints, and the overall centroid were selected for the experiments, for a total of seven concentration profiles for each of the 158 lipid combinations.
Thus, the total number of possible combinations is given by 75 (buffers) * 158 (qualitative lipid mixtures) * 7 (quantitative lipid possibilities) = 82,950 (total experimental space).
Applying conventional DoE methods to a complex system like this is very difficult. We attempted to generate a screening design using two major software packages (JMP and Design-Expert) and found them incapable of dealing with multilevel qualitative factors with the sorts of constraints specified for the lipid library.
Experimentation using iterated high-throughput screening and Evo-DoE begins with an initial generation selected from the whole experimental space. In standard methods such as genetic algorithms, this selection is random. More recently, low-discrepancy random sequence protocols have been used to avoid the gaps and overlapping points that frequently arise in the random approach, by biasing the random sampling toward unsampled regions of the experimental space. This is essentially a high-dimensional version of the stochastic sampling used in Kriging
Predictions of experiments were obtained from a neural network model (learned with back-propagation using
This work was conducted as part of the European Union integrated project PACE (EU-IST-FP6-FET-002035). Norman Packard and Gianluca Gazzola acknowledge helpful conversations with Irene Poli. The authors also thank the European Center for Living Technology.