Skip to main content
AI Infrastructure

AWS makes SageMaker HyperPod model caching generally available

AWS makes SageMaker HyperPod model caching generally available Image: Primary
Amazon Web Services said model caching for SageMaker Inference on HyperPod is now generally available in all regions where HyperPod is offered. The feature pre-loads model weights and inference-server container images onto cluster nodes before pods are scheduled, so pods read from local NVMe storage at roughly 7 GB/s instead of pulling weights and images over the network. AWS said benchmarks across 57-145 GB models showed about 60 percent faster scale-out with weights caching, while image caching removed over two minutes of cold image-pull time. AWS said the first cache population still requires a remote download, and weights caches are per-node, so NVMe consumption scales with node count.
Sources
In this story
Published by Tech & Business, a media brand covering technology and business. This story was sourced from Artificial Intelligence and reviewed by the T&B editorial agent team.
Back to Newswire
Keep reading
Full wire
Infrastructure AI
Infrastructure AI

1NCE launches benchmarking and unified AI access for IoT customers

1NCE announced an upgrade to its Insights service and the launch of 1NCE AI. Insights gives customers access to aggregated benchmark data from more than 30,000 deployments for comparisons of network usage, battery performance, har...

Infrastructure Products
Infrastructure Products

AWS extends Lambda recursive loop detection to Europe Sovereign Cloud

AWS said Lambda recursive loop detection is now supported for functions running in Europe Sovereign Cloud. The feature automatically detects and stops recursive invocations between Lambda functions and supported services such as A...

Products Infrastructure
Products Infrastructure

AWS Transform for .NET adds automatic unit test generation for modernized code

AWS said AWS Transform for .NET can now automatically generate unit tests for the code it modernizes. When enabled, the service targets testable classes in a transformed .NET application, such as business logic and controllers, pr...

AI Products
AI Products

Amazon Quick desktop AI assistant reaches general availability

Amazon says its Quick desktop application is now generally available on macOS and Windows, following a preview used by customers in manufacturing, healthcare and sports. The company claims Quick runs on AWS infrastructure custome...