Skip to main content
AI

AWS introduces Ray Serve container for model inference

AWS introduces Ray Serve container for model inference Image: Primary
AWS has introduced a Ray Serve Deep Learning Container for inference workloads, positioning it as a migration option for teams using unmaintained TorchServe. AWS says the image bundles PyTorch, Ray Serve, FastAPI, Uvicorn and GPU-stack dependencies in tested configurations, with separate images for EKS, EC2 and SageMaker. The company demonstrated the GPU image serving Qwen3-VL-2B on a single EKS g5.xlarge node with one NVIDIA A10G GPU. AWS says the containers receive security patches at build time.
Sources
In this story
Published by Tech & Business, a media brand covering technology and business. This story was sourced from AWS Machine Learning Blog and reviewed by the T&B editorial agent team.
Back to Newswire
Keep reading
Full wire
Capital AI
Capital AI

Cognition reports $2 billion Series E at $48 billion valuation

Cognition AI said it raised more than $2 billion in a Series E financing at a $48 billion valuation. SiliconANGLE reported that the company's prior May round raised more than $1 billion at a $26 billion valuation and that Cogniti...

AI
AI

AWS makes GPT-6 Astra available through Bedrock

Amazon Web Services said GPT-6 Astra is now generally available through Amazon Bedrock APIs. AWS said customers can configure ChatGPT Work and Codex to use the model on Bedrock, while new ChatGPT Work enterprise plugins extend bro...

AI
AI

SageMaker expands MLflow registry sync with metrics and lineage

AWS says its Managed MLflow integration with SageMaker AI Model Registry now transfers training metrics, evaluation results, inference specifications and lineage when registered models are synchronized. The opt-in sync also suppor...

AI Infrastructure
AI Infrastructure

OpenAI contracts Malaysia AI compute from Firmus

OpenAI has contracted dedicated AI compute from two Firmus AI Factory sites in Malaysia under a multi-year agreement announced September 8. Firmus said OpenAI becomes an anchor customer and that its total contracted capacity acros...

AI
AI

AWS adds feature-level writes to SageMaker Feature Store

AWS said SageMaker Feature Store now lets users update individual feature values within a record rather than rewrite the entire record. The company said a single call can replace read-modify-write workflows, leave unspecified feat...

AI
AI

AWS adds direct long-term memory ingestion to AgentCore

AWS has added an IngestData API to Amazon Bedrock AgentCore Memory that sends content directly to configured long-term memory extraction strategies. Previously, content had to be stored first as a short-term memory event before it...