Accelerating Drug Discovery by Combining Quantum-Based Models w/ Machine Learning | NVIDIA GTC 2024

NVIDIA Developer · Advanced ·📐 ML Fundamentals ·2y ago

Key Takeaways

The video demonstrates the acceleration of drug discovery by combining quantum-based models with machine learning, leveraging tools such as NV M library, Tinker HP, and Deep HP platform to achieve state-of-the-art performance in predicting binding affinity of small molecules to macromolecular targets. The Phoenix equivalent approach is also introduced, which combines machine learning and physics-based approaches to improve binding affinity results.

Full Transcript

[Music] hello I am Lua Garder from sban University and I will be showing you today how Quantum based models in the context of molecular simulations can accelerate drug Discovery and how machine learning approaches can be used in that framework to further accelerate this intrinsically complex problem first some words about molecular simulations and about their role in a drug Discovery context in molecular simulations we make systems of atoms move through an interaction po potential and the application of Newton's lows what that means in practice is that this interaction potential the model that we use to represent the interactions between atoms is crucial and the quality of the results of any simulation will depend on the accuracy of this model in the context of drug Discovery we want to find small molecules that binds to a macromolecular target of Interest this can be a protein or RNA for example and we need this binding to happen in such a way that they inhibit their function in practice we need to find a zone of interest in the Target a drugable pocket where such a small molecule could bind and then to model this binding process to do so because of the complexity and the dynamical nature of these systems we need to resort to highly accurate models to be able to be predictive we can draw a parallel between what I just described and the lending of a space probe to the irregular surface of an asteroid as represented in this slide finally a highly accurate model is Not Enough by itself because we also need to extensively sample confirmations of the target alone to find dragable pockets and the relative movements of the Target and the smaller molecule to gain significant insight about their binding that means that we need high performance implementations of these models to be able to use them in a practical drug Discovery framework now some informations about the new generation molecular models that we have been working on we know that the ground truth would require a Quantum description of all the atoms and electrons of the molecules we are dealing with but for Meaningful systems it is way too expensive computationally so we need to make some approximations traditionally for the last decades people have used two B pairwise models to represent all the interactions and especially the electrostatic ones but in practice when need to include anisotropy and also some kind of response to a change in the environment which is typically what Quantum mechanic mechanics would yield to do so we include a multipolar description of the density of charge of our molecules this gives us an isotropy and we include many body polarization to represent the electronic mobility of our system when it changes of conformation so what would these more accurate models give us in terms of properties of interest here are a list of things that would become possible in this context first we can deal with small subtle energetical difference of the complex bios systems because of our accurate description of all and potentially weak interactions then this gives us access to the computation of free energies of binding useful for drug design and I will come back to that a little bit later also it means that we can model Metals including heavy ones and also ionic liquids that play a pivotal role in modern batteries because of the inclusion of many body polarization effects also and I won't develop that too much but just note that when coupled to a purely qm description of a subers of the system this new generation of models yield accurate liquid phase spectroscopy and reactivity of complex Spees remember that in practice we need massive sampling and not only high resolution this is where high performance Computing comes into play and this has been our Focus for the past years to give you some historical context 10 years ago no real high performance implementation of polarizable models existed and the most popular code handling this thinker was only accelerated through a shared memory par parallelism through openmp directives which was not enough to tackle real life problems then we spent a lot of time and energy to start from this popular and well-designed code thinker to develop a high performance version of it that we call Tinker HP but because of the complexity of the algorithms involved in the resolution of this more complex equations we had to rethink fundamental aspects of the code in its structure and we had to design specific approaches to be able to reach scalable simulations these are described in this chemical science paper from 2018 in practice we started with an MPI based CPU implementation in Fortran in double precision being able to scale on tens of thousands of course on large enough systems but also being able to scale on smaller clusters such as the ones typically present in academic Labs so we were petascale ready but with the Advent of modern gpus especially the ones from Nvidia we knew that we could gain considerable performance by leveraging these platforms this is why we developed a specific version of our Tinker HP code dedicated to the use of potentially multiple Nvidia gpus this work was mainly done by Olivia AA who also works at Subban here also specific algorithmic developments and adjustments had to be made in order to fit the specificities of gpus with a with a special care given to Precision all the more since we need to solve the mbody polarization equations through an iterative procedure that is sensitive to Precision we published a paper describing all of these specifics in the Journal of chemical Theory and computation in 2021 okay just a few words about the GPU implementation itself we have two different GPU implementations the first and the one we started with uh completely relies on an open ACC Portage open ACC directives are used to both transfer the data between the CPU host and the GPU device and also to run the computationally Intensive Kels on Nvidia gpus we use this implementation to run double Precision simulations taking advantage of both V100 and a100 HPC cards the second one still resorts to open ACC directives to handle data transfer between the GPU and the host but uses heavily optimized Cuda kernels to further gain performance this implement ation is used in what is called in the context of molecular Dynamics mixed Precision mode that is to say that the energy and the force kernels that are the most computationally intensive use single Precision but that the data is then accumulated in double Precision this yielding an optimal balance between performance and precision a key feature of both these implementations is that almost all operations are uploaded on the gpus this limiting synchron between the CPU and the GPU furthermore the multiple GPU implementation follows the same 3D domain decomposition logic of the cpu1 and we make sure that the data transfer happen directly between gpus through Cuda aware MPI library now some practical benchmarks with systems of Representative sizes from the well-known dhfr Benchmark that is made of a little bit more than 20,000 atoms to larger protein systems such as such as the main protees of sarov 2 in solution which contains around 100,000 atoms and up to much larger multi-million atom systems such as what we have called the C system that contains more than 7 million atoms all these benchmarks have been run on the Jean super computer of of geni in France and on the Celine supercomputer of Nvidia the cin super computer of Nvidia is made of dgx a100 leag together in practice within each of these dgx a100 eight a100 cars are efficiently linked together through the EnV switch technology to further improve memory distribution and minimize data exchange between the gpus we resort to the NV M library all in all we observe the best performance ever obtained both on gpus and CPUs on all of these systems for these kind of models that is to say that we have more than 40 NS per day of production for the D HFR system more than 4 nond per day on the well-known stmv system uh that is made of more than 1 million atoms and almost one nond per day for the C system that I mentioned a little bit before what we see is that for the smaller systems there is no real gain of using more than one GPU which is expected given the amount of computer of compute power that one GPU alone already contains but by growing the system size we see that starting from systems of around 200,000 atoms there is an interest of using several gpus and then there is a significant gain for the larger systems this gain depending on the system size is more clearly Illustrated in this figure that shows in a logarithmic scale the performance gain from one to a gpus on the cind superc computer for all of these systems to some things up about Tinker HP and its efficient multi-gpu implementation we see that for the stmv system that is represented here and that is made as I mention before of around 1 million atoms we reached an acceleration of a factor around 6,000 from the original Tinker open MP implementation on eight CPU CES to the multiple GPU Cuda implementation it means that it has dramatically opened the field of potential applications of the new generation polarizable models and this is what I'm going to talk about in the rest of this presentation one of the main application that I want to focus on now is the prediction of binding Affinity of small molecues to a macromolecular Target in a drug Discovery project a well- defined Target for example a protein is chosen and then the goal is to find a small molecule that binds to it in such a manner that it Alters its function typically Pharmaceuticals companies rely on high throughput screening of already known compounds to get initial Heats molecules and then on a large number of synthesis to optimize such hits this is why being able to reliably predict the actual binding Affinity of small molecules to targets is extremely appealing in the context of drug Discovery because numerical simulations could drastically reduce the number of synthesis of drunk candidates candidates to be made and thus reduce the time and money spent on such projects but what does that mean in practice in practice we want to compute the free energy difference between the bound state where the liant or small molecule lies in The Binding pocket of the host as represented here in this slide and the Unbound state where both the small molecule and the host are in the bulk in solution without interacting between each other rather than sampling explicitly this binding and unbinding process what is routinely used nowadays uh are the so-call alchemical methods where the small molecule is first gradually decoupled from the host by scaling down progressively its interactions with it and then separately we compute the free energy of decoupling the lion from the bulk in a similar fashion and finally because free energy is a state function we can then recover the free energy of interest from the thermodynamical cycle as represented here in this slide without going into the details of how simulations are run to recover these quantities in practice let's just say that it requires numerous and long molecular Dynamics trajectories hence the key importance of a high performance code in practice many benchmarks of represent ative host gas systems have been run in the past few years through the blind sample challenges where academic groups try to reproduce blindly experimental binding affinities of small molecules to larger host the polarizable and multipolar amiba model has repeatedly perform really well on these challenges as you can see on the slide here that represents the results of the sample challenge of last year it concerns The Binding of several small molecules to the host system that is shown here on the left on the right you can can see the correlation between the computed binding affinities with this highly accurate model and the experimental ones and you can also see that they are in agreement within 1 kilal per mole for almost all of them which constitutes state-of-the-art performance so this shows how the high performance implementation of new generation models can help in a practice in drug in a drug Discovery project now let's go back to the model itself as we saw that it drives the quality of all the molecular simulations what we know is that even with the advanced polarizable models we still lack some key Quantum effects that happen at short range another limitation of the models that I presented before is that they are by Nature non-reactive we also know that there are now well-designed and well curated data sets of quantum mechanical data that enable the development of machine learning interaction Potentials in which all the interatomic interactions are described with machine learning methods without going into the details let me just mention that in general such models L leverage a short range descriptor of the local environment of each atom which is then given to a neural network to predict the energies what that means in practice is that these models are in general slower than polarizable force fields but as they are intrinsically short range they are scalable as you can see here on this slide these models are implemented in Tinker HP through the Deep HP platform that is described in a chemical science paper that we published last year okay so these models perform well in many cases but by design they like in general long range effects for which the physics is well known for example we know that electrostatics at long range follows Kon law this led us to looking at ways to combine machine learning approaches at short range and physics based ones at long range and for that we followed two different routs the first idea is to use a nonbing scheme similar as what is done in mixed quantum mechanical molecular mechanics methods where the interactions within the system of Interest are treated with the machine learning potential and the rest of the inter actions are treated with the polarizable force field the second IDE is more complex we use machine learning models to predict Atomic properties such as charges volumes and so on that will thus depend on the environment and that we use within a classical functional form with coolant potential FAL interactions and so on this is what we have called The Phoenix equivalent approach Phoenix stands for force field enhanced Nar Network interactions and we published this recently in chemical science one of one really nice feature of is that it is reactive now let's look at practical results that we obtain with both approaches on this slide you can see binding Affinity results obtained with both the amiba polarizable force field alone in yellow and by combining the Ani machine learning model an a well known machine learning model with aiar through the edding I just described we see that by including the machine learning model at short range we manage to improve the overall results especially for some systems we go in practice from a root me Square ER of 1.81 kilal per mole to one of a little bit less than 1 kilal per mole so we reduce the error by around the factor two and finally one results that illustrates the reactive nature of the Phoenix approach you can see here several dissociation curves where we make one atom of a molecule leaves its equilibrium position here we have three molecules we have CH4 water and HF we can see that the stateof the art any 2x model by itself gives unphysical results whereas the Phoenix ones are much closer to reference values that are shown in the doted curves we are currently actively working on enriching the Phenix models by incorporating additional energy temps to it such as explicit polarization and charge transfer the Phoenix approach is really promising as it naturally leverages the strength of machine learning methods to handle the complex interactions at short range and the robustness of the physics based long range interactions all in all I am convinced that with the ever increasing compute power and the always richer data sets of qm data pH CL will bring critical additional Insight in drug Discovery for example it will allow us to tackle efficiently the question of calent Inhibitors thanks to its reactive nature to conclude I want to thank all the people and institutions that have contributed to the results that I have been showing today there are people of course from San University but also from the US and from a company that I have confounded with people from France and the United States which is called Cubit Pharmaceuticals and that leverage routinely all of these tools to find new drugs I thank you all for your [Music] [Music] attention [Music] [Applause] y [Music] [Music] [Applause] w [Applause] [Music] [Music] w [Music] [Music] [Applause] [Music] [Applause] [Music] oh [Applause] [Music] [Applause] [Music] [Applause] [Music] oh [Music] [Applause] [Music] [Applause] [Music] [Music] [Applause] [Music] w [Music] [Applause] [Music] [Applause] [Music] [Applause] [Music] w [Music] [Applause] [Music] [Applause] [Music] [Applause] w [Music] [Applause] [Music] [Applause] [Music] w [Music] yeah

Original Description

The first stages of drug discovery involve finding a molecule with a good affinity to a protein target of interest. It's a long and costly process with a low success rate, but it can be drastically accelerated by in silico molecular simulations, provided that these are accurate and fast enough. This session presents key advances in this direction that leverage a unique combination of quantum-based approaches with machine learning in a massively multi-GPU context. Speaker: Louis Lagardère, Research Engineer and Co-Founder, Sorbonne Université and Qubit-Pharmaceuticals Explore more GTC 2024 sessions like this on NVIDIA On-Demand: https://nvda.ws/3U33qo7 Read and subscribe to the NVIDIA Technical Blog: https://nvda.ws/3XHae9F Original GTC 2024 Session: Combining Quantum-Based Models With Machine Learning Accelerates Drug Discovery [S61502] #GTC24 #NVIDIA #GTC #AI #DrugDiscovery #QuantumComputing #Simulation #Modeling #LifeSciences #Pharma
Watch on YouTube ↗ (saves to browser)
Sign in to unlock AI tutor explanation · ⚡30

Playlist

Uploads from NVIDIA Developer · NVIDIA Developer · 0 of 60

← Previous Next →
1 Ray Tracing Essentials Part 2: Rasterization versus Ray Tracing
Ray Tracing Essentials Part 2: Rasterization versus Ray Tracing
NVIDIA Developer
2 Ray Tracing Essentials Part 3: Ray Tracing Hardware
Ray Tracing Essentials Part 3: Ray Tracing Hardware
NVIDIA Developer
3 Ray Tracing Essentials Part 4: The Ray Tracing Pipeline
Ray Tracing Essentials Part 4: The Ray Tracing Pipeline
NVIDIA Developer
4 NsightGraphics 2020 2 Release Spotlight
NsightGraphics 2020 2 Release Spotlight
NVIDIA Developer
5 Ray Tracing Essentials Part 5: Ray Tracing Effects
Ray Tracing Essentials Part 5: Ray Tracing Effects
NVIDIA Developer
6 Ray Tracing Essentials Part 6: The Rendering Equation
Ray Tracing Essentials Part 6: The Rendering Equation
NVIDIA Developer
7 Ray Tracing Essentials Part 7: Denoising for Ray Tracing
Ray Tracing Essentials Part 7: Denoising for Ray Tracing
NVIDIA Developer
8 Spatiotemporal Importance Resampling for Many-Light Ray Tracing (ReSTIR)
Spatiotemporal Importance Resampling for Many-Light Ray Tracing (ReSTIR)
NVIDIA Developer
9 Announcing Cloud-Native Support for Jetson Platform
Announcing Cloud-Native Support for Jetson Platform
NVIDIA Developer
10 JetsonTV: Build your next project with NVIDIA Jetson
JetsonTV: Build your next project with NVIDIA Jetson
NVIDIA Developer
11 Nsight Compute Feature Spotlight: Roofline Analysis, Asynchronous Copy, Sparse Data Compression
Nsight Compute Feature Spotlight: Roofline Analysis, Asynchronous Copy, Sparse Data Compression
NVIDIA Developer
12 Nsight Systems Feature Spotlight: OpenMP
Nsight Systems Feature Spotlight: OpenMP
NVIDIA Developer
13 Isaac Sim 2020: Deep Dive
Isaac Sim 2020: Deep Dive
NVIDIA Developer
14 NVIDIA Jetson: Enabling AI-Powered Autonomous Machines at Scale
NVIDIA Jetson: Enabling AI-Powered Autonomous Machines at Scale
NVIDIA Developer
15 NVIDIA Tools to Train, Build, and Deploy Intelligent Vision Applications at the Edge
NVIDIA Tools to Train, Build, and Deploy Intelligent Vision Applications at the Edge
NVIDIA Developer
16 Jetson Xavier NX Developer Kit: The Next Leap in Edge Computing
Jetson Xavier NX Developer Kit: The Next Leap in Edge Computing
NVIDIA Developer
17 Synthesizing High-Resolution Images with StyleGAN2
Synthesizing High-Resolution Images with StyleGAN2
NVIDIA Developer
18 NVIDIA Robotics: Isaac SDK and Sim 2020.1
NVIDIA Robotics: Isaac SDK and Sim 2020.1
NVIDIA Developer
19 Accelerating COVID-19 Research with GPUs
Accelerating COVID-19 Research with GPUs
NVIDIA Developer
20 Visualizing 150 Terabytes of Data
Visualizing 150 Terabytes of Data
NVIDIA Developer
21 Boosting Performance and Utilization with Multi-Instance GPU
Boosting Performance and Utilization with Multi-Instance GPU
NVIDIA Developer
22 Running Multiple Workloads on a Single A100 GPU
Running Multiple Workloads on a Single A100 GPU
NVIDIA Developer
23 NVIDIA Nsight Feature Spotlight: GPU Trace
NVIDIA Nsight Feature Spotlight: GPU Trace
NVIDIA Developer
24 Spark 3 Demo: Comparing Performance of GPUs vs. CPUs
Spark 3 Demo: Comparing Performance of GPUs vs. CPUs
NVIDIA Developer
25 NVIDIA Jetson Nano Wins Edge AI and Vision Alliance Award
NVIDIA Jetson Nano Wins Edge AI and Vision Alliance Award
NVIDIA Developer
26 NVIDIA IndeX on Google Cloud Platform Marketplace
NVIDIA IndeX on Google Cloud Platform Marketplace
NVIDIA Developer
27 DeepStream SDK: Best practices for performance optimization
DeepStream SDK: Best practices for performance optimization
NVIDIA Developer
28 Efficiently Deploying GPU Accelerated 5G CloudRAN for Edge AI Inferencing
Efficiently Deploying GPU Accelerated 5G CloudRAN for Edge AI Inferencing
NVIDIA Developer
29 NVIDIA PhysicsNeMo - Accelerating Scientific & Engineering Simulation Workflows with AI
NVIDIA PhysicsNeMo - Accelerating Scientific & Engineering Simulation Workflows with AI
NVIDIA Developer
30 NVIDIA Deep Learning Institute Instructor-Led Training Available Remotely
NVIDIA Deep Learning Institute Instructor-Led Training Available Remotely
NVIDIA Developer
31 Advancing AR Glasses
Advancing AR Glasses
NVIDIA Developer
32 Blender Cycles: RTX On
Blender Cycles: RTX On
NVIDIA Developer
33 Real-Time GPU-Accelerated Data Analytics of 250 million Flight Data Records of 737 Max grounding
Real-Time GPU-Accelerated Data Analytics of 250 million Flight Data Records of 737 Max grounding
NVIDIA Developer
34 Assessing Property Damage with AI
Assessing Property Damage with AI
NVIDIA Developer
35 RAPIDS: GPU-Accelerated Data Analytics & Machine Learning
RAPIDS: GPU-Accelerated Data Analytics & Machine Learning
NVIDIA Developer
36 DaVinci Resolve Turns RTX On
DaVinci Resolve Turns RTX On
NVIDIA Developer
37 RAPIDS with Plotly Dash : GPU-Accelerated Census 2010 Visualization
RAPIDS with Plotly Dash : GPU-Accelerated Census 2010 Visualization
NVIDIA Developer
38 NVIDIA IndeX for arivis5D Cloud Platform
NVIDIA IndeX for arivis5D Cloud Platform
NVIDIA Developer
39 NVIDIA Backchannel: Behind the Scenes of Marbles at Night RTX
NVIDIA Backchannel: Behind the Scenes of Marbles at Night RTX
NVIDIA Developer
40 NVIDIA Backchannel: Sneak Peek into Marbles RTX in Omniverse
NVIDIA Backchannel: Sneak Peek into Marbles RTX in Omniverse
NVIDIA Developer
41 How to Create "Paint" in Substance Painter
How to Create "Paint" in Substance Painter
NVIDIA Developer
42 Accelerate AI development for Computer Vision on the NVIDIA Jetson with alwaysAI
Accelerate AI development for Computer Vision on the NVIDIA Jetson with alwaysAI
NVIDIA Developer
43 Securing Next Generation Apps over VMware Cloud Foundation with Bluefield-2 DPU
Securing Next Generation Apps over VMware Cloud Foundation with Bluefield-2 DPU
NVIDIA Developer
44 Accelerated Data Centers with NVIDIA and VMware
Accelerated Data Centers with NVIDIA and VMware
NVIDIA Developer
45 GPU-Accelerated Motion Blur in Blender Cycles
GPU-Accelerated Motion Blur in Blender Cycles
NVIDIA Developer
46 NVIDIA Clara Guardian Virtual Patient Assistant
NVIDIA Clara Guardian Virtual Patient Assistant
NVIDIA Developer
47 Revolutionizing Supercomputing with NVIDIA UFM Cyber-AI
Revolutionizing Supercomputing with NVIDIA UFM Cyber-AI
NVIDIA Developer
48 Inventing Virtual Meetings of Tomorrow with NVIDIA AI Research
Inventing Virtual Meetings of Tomorrow with NVIDIA AI Research
NVIDIA Developer
49 Learning a Contact-Adaptive Controller for Robust, Efficient Legged Locomotion
Learning a Contact-Adaptive Controller for Robust, Efficient Legged Locomotion
NVIDIA Developer
50 Getting started with Jetson Nano 2GB Developer Kit
Getting started with Jetson Nano 2GB Developer Kit
NVIDIA Developer
51 NVIDIA Jetson Developer Community AI Projects
NVIDIA Jetson Developer Community AI Projects
NVIDIA Developer
52 Open-source projects on NVIDIA Jetson Nano 2GB Developer Kit
Open-source projects on NVIDIA Jetson Nano 2GB Developer Kit
NVIDIA Developer
53 Real-Time Ray Tracing with Project Lavina
Real-Time Ray Tracing with Project Lavina
NVIDIA Developer
54 Jetson AI Fundamentals - S1E2 - Hello Camera
Jetson AI Fundamentals - S1E2 - Hello Camera
NVIDIA Developer
55 Develop Optimized Conversational AI Models with NVIDIA NeMo on DGX A100
Develop Optimized Conversational AI Models with NVIDIA NeMo on DGX A100
NVIDIA Developer
56 Jetson AI Fundamentals - S1E4 - Image Regression Project
Jetson AI Fundamentals - S1E4 - Image Regression Project
NVIDIA Developer
57 Jetson AI Fundamentals - S2E1 - JetBot Intro and Hardware
Jetson AI Fundamentals - S2E1 - JetBot Intro and Hardware
NVIDIA Developer
58 Jetson AI Fundamentals - S2E2 - JetBot Software Setup
Jetson AI Fundamentals - S2E2 - JetBot Software Setup
NVIDIA Developer
59 Jetson AI Fundamentals - S1E1 - First Time Setup with JetPack
Jetson AI Fundamentals - S1E1 - First Time Setup with JetPack
NVIDIA Developer
60 Jetson AI Fundamentals - S1E3 - Image Classification Project
Jetson AI Fundamentals - S1E3 - Image Classification Project
NVIDIA Developer

The video teaches how to accelerate drug discovery by combining quantum-based models with machine learning, using tools such as NV M library and Deep HP platform. The Phoenix equivalent approach is introduced to improve binding affinity results. By applying machine learning models to molecular dynamics data, researchers can improve the accuracy of binding affinity predictions and accelerate the drug discovery process.

Key Takeaways
  1. Implement quantum-based models using NV M library
  2. Train machine learning models on molecular dynamics data using Deep HP platform
  3. Combine machine learning and physics-based approaches using Phoenix equivalent approach
  4. Evaluate the performance of machine learning models on binding affinity predictions
  5. Refine machine learning models to improve accuracy
💡 The combination of quantum-based models and machine learning can significantly improve the accuracy of binding affinity predictions and accelerate the drug discovery process.

Related Reads

Up next
Is coding becoming obsolete? | Find out what's the new fundamentals
SCALER
Watch →