AWS Podcast cover image

#600: Amazon SageMaker Multi Model Endpoints

AWS Podcast

00:00

Multi-Model Endpoints Support for GPU

In MME for each endpoint that customers provision, we have a shared free defense instances behind. And then customer tells us, hey, I want this instance type, I want to provision, let's say, 10 instances for it. And then what MME tries to do is you'll also be telling us the models where they are in, right? So MME loads and unloads models on the shared fleet of instances that you have provisioned based on the traffic that we see. Then it tries to optimize for the cost by understanding your traffic pattern.

Transcript
Play full episode

Remember Everything You Learn from Podcasts

Save insights instantly, chat with episodes, and build lasting knowledge - all powered by AI.
App store bannerPlay store banner