← Back

models

Introducing cross-Region inference for OpenAI GPT-5.6 models on Amazon Bedrock

Amazon Bedrock now supports OpenAI GPT-5.6 models in over 25 AWS Regions with cross-Region inference, enabling scalable and flexible deployment options for AI workloads.

AS1 NewsSource: aws.amazon.com

modelsawsgpt-5-6
OpenAI$1,487.99-1.17%AMZN$256.78-0.82%REAL$0.0751+2.65%

Amazon Bedrock has expanded its support for OpenAI GPT-5.6 models, including variants Sol, Terra, and Luna, across more than 25 AWS Regions. This update introduces cross-Region inference (CRIS), allowing requests to be routed between Regions based on capacity and geographic requirements. The models accept text and image inputs, support reasoning, server-side tool calling, and prompt caching, and can be accessed via OpenAI and Converse APIs.

CRIS works through inference profiles—either geographic, which restrict routing within a specific geography, or global, which routes across all supported Regions based on real-time capacity. This mechanism enhances throughput and maintains performance under load, especially for workloads with high demand.

Users can invoke these models through the Amazon Bedrock console's text playground or programmatically via APIs, including the OpenAI Responses API, Chat Completions API, and the Amazon Bedrock Converse API. The models support streaming responses, making them suitable for real-time applications.

Security and compliance are maintained through AWS IAM policies, VPC endpoints, and CloudTrail logging. Data processed through CRIS may cross Regions, but all requests are authenticated and authorized under the same security model as in-region calls.

For workload management, quotas are tracked per inference profile, with separate allocations for geographic and global profiles. Monitoring tools like Amazon CloudWatch provide real-time metrics on usage, latency, and errors, aiding in capacity planning.

This deployment enhances the flexibility and scalability of AI applications on AWS, enabling organizations to leverage high-capacity, low-latency inference across multiple Regions while adhering to data residency requirements.

neutral

Enables scalable deployment of GPT-5.6 models with cross-Region inference, improving throughput and performance for AI workloads.