Developer
News and Updates
Get Support
Sign in
Get Support
Sign in
DOCUMENTATION
Cloud
Data Center
Resources
Sign in
Sign in
DOCUMENTATION
Cloud
Data Center
Resources
Sign in
Last updated Sep 30, 2026

Forge LLM limits

The following limits apply for each installation of your app when using the Forge LLMs API:

ResourceLimitDescription
Context window size in tokens (in / out)Varies by model. See Context window limits by model.The maximum number of input tokens a model can reference, and the maximum number of output tokens it can subsequently generate.
Requests per minute100The number of prompts sent to any model in any given minute.
Tokens per minute500,000The maximum number of tokens that a single model can process each minute.
Inference time in minutes5The maximum time a model can process and generate responses before a timeout occurs, assuming the Async events API is used with a specified timeout equal or greater than 5 minutes. Otherwise the specified or default timeouts apply.

Context window limits by model

The context window size depends on the model tier you use. For the full list of supported models and their tiers, see Forge LLMs models.

Model tierContext window (in / out)
Haiku200K / 64K
Sonnet1M / 128K
Opus1M / 128K

Rate this page: