Skip to content

llm_gemini: the custom profile's Output Tokens field saves under a key nothing reads, and no output limit reaches Google #2529

Description

@krishhgg

Description

The Gemini node's Custom profile has an Output Tokens field. Whatever you type into it has no effect. Two separate problems cause this.

First, the field saves its value under the wrong key. Its id is gemini.outputTokens, and the engine names a field after the last part of its id, so the value lands in the config as outputTokens. ChatBase, the base class every LLM node shares, reads modelOutputTokens. Every other LLM node uses that name. ChatBase therefore ignores the value and uses the 65,536 that the Custom profile carries in services.json.

Second, the Gemini driver never sends an output limit to Google. It calls generate_content with only the model name and the prompt. So even a correctly named field would only change ChatBase's own numbers: the debug log, the getOutputLength answer, and the chunk size the LLM preprocessor derives from it. Google would still apply its own default.

The config code already has a warning for this exact mistake. Config._suggestKey answers outputTokens with "did you mean modelOutputTokens?" (#2152, fixed in #2166). It stays silent for Gemini. The list of known keys is built partly from the node's own field names, and this node declares gemini.outputTokens, so outputTokens counts as a real key.

The repo ships a pipeline with this key. pipelines/git_agent_example.pipe sets "outputTokens": 8192 on its custom Gemini node, and that 8,192 is ignored today. #2152 mentions an older pipeline that "carried outputTokens in every llm_gemini node"; the field itself is the likely source.

Steps to Reproduce

  1. Load pipelines/git_agent_example.pipe, or add a Gemini node, pick the Custom profile, set a model, and set Output Tokens to 8192.
  2. Run the pipeline with debug logging on. ChatBase logs Output tokens : 65536, not 8192.
  3. Ask for a long answer. The request to Google carries no max_output_tokens, so the answer can run past 8,192 tokens.

Without a running engine, resolving the example's custom block through the real Config.getNodeConfig (with getServiceDefinition stubbed to return the node's services.json) gives:

{'model': 'gemini-3.1-flash-lite-preview', 'modelTotalTokens': 1000000, 'modelOutputTokens': 65536, 'outputTokens': 8192}

No warning is logged.

Expected vs Actual Behavior

A value typed into Output Tokens should become the node's output limit. ChatBase should use it, and the request to Google should carry it as max_output_tokens.

Today the value sits in the config as outputTokens, ChatBase uses 65,536, and Google receives no limit at all.

Severity

P3 (low, minor inconvenience). Answers are not cut short. The cost is that nobody can cap Gemini output from the node, and the preprocessor sizes chunks from a number the user did not choose.

Platform

All. The code is Python and not platform specific.

Version / Environment

develop at da5ac85.

What a fix needs to do

  • Rename the field so it saves as modelOutputTokens. Use gemini.modelOutputTokens, or the shared modelOutputTokens field id that llm_baidu_qianfan uses.
  • Send the limit to Google in Chat._chat, as max_output_tokens in the config argument of generate_content.
  • Decide what happens to pipelines that already saved outputTokens, and update the example pipeline in the same change.

Pipelines that saved the old key

Renaming loses no working value, because the old value never worked. What changes is what people see.

Without an alias, a saved outputTokens stays ignored and the form shows an empty Output Tokens field, so the number the user typed disappears from view. After the rename, outputTokens is no longer a declared field, so the misnamed-key warning starts to fire for these pipelines. That warning says the key "is ignored", which would be accurate.

With an alias (read outputTokens when modelOutputTokens is absent), old values start to take effect. That is a behavior change. ChatBase refuses to start with a value under 1,024, and any value starts capping output once the driver sends it. The warning would then be wrong, because the key is no longer ignored.

Either choice is defensible, as long as it is made on purpose and tested.

How to test it

Resolve the Custom profile with a modelOutputTokens value and check that getOutputTokens() returns it. Then call the driver with a fake google.genai client and check that generate_content received the limit. Details below.

Related


For your agents

Dense context for an AI agent picking this up. Line numbers are from develop at da5ac85; re-check them before editing. Items marked [verify] were inferred from reading code, not observed.

Where the value goes

Why no limit reaches Google

  • Chat._chat at nodes/src/nodes/llm_gemini/gemini.py:194 calls self._client.models.generate_content(model=self._model, contents=prompt) with no config.
  • The node sets no _llm and no _native_stream_provider, so ChatBase.chat_string always ends in _chat_with_retries, which calls _chat. There is no other request path to patch.
  • The save-time probe in IGlobal.validateConfig also calls generate_content without a config. Leave it alone; it sends "Hi".

Why the warning stays silent

  • Config._knownConfigKeys at packages/ai/src/ai/common/config.py:90 adds every field name with its prefix stripped, so gemini.outputTokens contributes outputTokens.
  • Config._suggestKey at config.py:109 returns None at once when the key is known (lines 127 to 128). Its docstring names this exact dropped-prefix case.
  • Confirmed by running the real config.py against the node's services.json: 'outputTokens' in knownKeys is True and _suggestKey('outputTokens', knownKeys) is None.

Files that carry the old name

What a fix must do

  1. Rename the field id. gemini.modelOutputTokens keeps the node's naming; the bare modelOutputTokens matches llm_baidu_qianfan (services.json:135). Update the Custom form's properties list too.
  2. Pass the limit in _chat: config=types.GenerateContentConfig(max_output_tokens=self._modelOutputTokens) from google.genai.types, or the equivalent dict. Keep _report_gemini_usage and the GeminiNoTextError handling as they are.
  3. Choose alias or no alias for outputTokens (see the human section). If you alias, do it in gemini.py after super().__init__, run the value through validate_max_tokens, and keep the 1,024 floor in mind.
  4. Move the example pipeline to modelOutputTokens.

Traps

  • Sending a limit changes every catalogue profile, not only Custom. Each will start sending its modelOutputTokens (65,536, 65,535, 58,982, 32,768 or 8,192 in today's file). A value above what Google accepts for that model fails the call [verify per model]. The values come from the model sync (// openrouter and // sync_models.config.json comments), so a wrong one belongs in tools/sync_models/src/sync_models.config.json, not in a hand edit.
  • On Gemini 2.5 and later, thinking tokens count against max_output_tokens [verify]. A low user value can then end the response before any text, which surfaces as GeminiNoTextError with finish_reason=MAX_TOKENS. That error is already non-retryable, which is correct, but the message should stay readable.
  • validate_max_tokens clamps the output limit to modelTotalTokens. A custom model with a small context window gets a small output limit, which is right, but test it.
  • After the rename, Config._warnMisnamedKeys will start warning on pipelines that still say outputTokens. If you add an alias, either suppress that warning for this key on this node or reword it, because "is ignored" would be false.
  • The stub in test_gemini_token_metrics.py provides only google.genai.Client. If gemini.py starts importing google.genai.types, that stub needs a types attribute too, or pass a plain dict as config.
  • nodes/test/mocks/google/genai/__init__.py:33 already accepts config=None in generate_content, so the node's mock test keeps passing whether or not a config is sent. It will not catch a regression unless a test asserts on the argument.

How to test

  • Add a unit test in nodes/test/llm_gemini/. The existing test_gemini_token_metrics.py shows how to load the node module with a fake genai. Build Chat with {'profile': 'custom', 'custom': {'model': 'm', 'apikey': 'k', 'modelOutputTokens': 8192}} and assert getOutputTokens() == 8192.
  • Use the same setup with a fake client that records its arguments. Call _chat('hi') and assert config.max_output_tokens == 8192.
  • Check that a catalogue profile sends its catalogue value.
  • If you alias, test that outputTokens: 8192 alone gives 8192, modelOutputTokens wins when both are set, and outputTokens: 512 fails with ChatBase's 1,024 message.
  • Resolve pipelines/git_agent_example.pipe after the change and check the 8,192 reaches ChatBase.

Acceptance criteria

  • The Custom form's Output Tokens value is what ChatBase reports and what Google receives as max_output_tokens.
  • outputTokens no longer appears in services.json, the README schema table, or the example pipeline.
  • The chosen behavior for old outputTokens values is tested and stated in the PR description.

Out of scope

  • llm_vision_gemini and accessibility_describe, which build their own generate_content calls.
  • Custom forms on other nodes that have no output field at all (Anthropic, OpenAI, xAI, Bedrock, Ollama and others). That gap is real but separate.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't workingmodule:nodesPython pipeline nodes

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions