Skip to main content
AI Products

AWS adds model caching to SageMaker HyperPod, cutting inference cold starts

Amazon Web Services said SageMaker HyperPod now supports model caching, an inference optimization that pre-loads model weights and container images onto cluster nodes so pods start in seconds rather than minutes. The feature is generally available in all regions where HyperPod is available. AWS describes two independent capabilities: a weights cache that stores model weights on local NVMe instead of pulling from S3 or FSx over the network, and an image cache that pre-pulls container images to skip ECR downloads. Pods landing on a node without a warm cache fall back to the original source automatically. AWS reports benchmarks across models from 57 GB to 145 GB showing roughly 60% faster scale-out, with the image cache cutting over two minutes of image-pull time, a 97% reduction. Customers enable it through the HyperPod Inference Operator.
Sources
In this story
Published by Tech & Business, a media brand covering technology and business. This story was sourced from Recent Announcements and reviewed by the T&B editorial agent team.
Back to Newswire
Keep reading
Full wire
AI Science
AI Science

NASA and IBM release open lunar foundation model on Hugging Face

NASA and IBM Research released the NASA-IBM Lunar Foundation Model, an open-source AI model built for lunar science, hosted publicly on Hugging Face with its codebase on GitHub. NASA said the model was trained primarily on 17 yea...

Security Policy
Security Policy

Florida confirms DMV driver database breach via stolen police credentials

The Florida Department of Highway Safety and Motor Vehicles confirmed that its DAVID driver database was breached after the ShinyHunters extortion gang claimed to have compromised the system. The agency said it learned of the bre...

Security
Security

Wiz reports Artifactory flaw chain exploited to plant Rust backdoor

Wiz says multiple threat actors chained two JFrog Artifactory vulnerabilities, CVE-2026-42018 and CVE-2026-42016, against self-hosted servers between August 15 and September 8, 2026, obtaining an internal anonymous-user JWT and ex...