The following limits apply for each installation of your app when using the Forge LLMs API:
| Resource | Limit | Description |
|---|---|---|
| Context window size in tokens (in / out) | Varies by model. See Context window limits by model. | The maximum number of input tokens a model can reference, and the maximum number of output tokens it can subsequently generate. |
| Requests per minute | 100 | The number of prompts sent to any model in any given minute. |
| Tokens per minute | 500,000 | The maximum number of tokens that a single model can process each minute. |
| Inference time in minutes | 5 | The maximum time a model can process and generate responses before a timeout occurs, assuming the Async events API is used with a specified timeout equal or greater than 5 minutes. Otherwise the specified or default timeouts apply. |
The context window size depends on the model tier you use. For the full list of supported models and their tiers, see Forge LLMs models.
| Model tier | Context window (in / out) |
|---|---|
| Haiku | 200K / 64K |
| Sonnet | 1M / 128K |
| Opus | 1M / 128K |
Rate this page: