Description
The Gemini node's Custom profile has an Output Tokens field. Whatever you type into it has no effect. Two separate problems cause this.
First, the field saves its value under the wrong key. Its id is gemini.outputTokens, and the engine names a field after the last part of its id, so the value lands in the config as outputTokens. ChatBase, the base class every LLM node shares, reads modelOutputTokens. Every other LLM node uses that name. ChatBase therefore ignores the value and uses the 65,536 that the Custom profile carries in services.json.
Second, the Gemini driver never sends an output limit to Google. It calls generate_content with only the model name and the prompt. So even a correctly named field would only change ChatBase's own numbers: the debug log, the getOutputLength answer, and the chunk size the LLM preprocessor derives from it. Google would still apply its own default.
The config code already has a warning for this exact mistake. Config._suggestKey answers outputTokens with "did you mean modelOutputTokens?" (#2152, fixed in #2166). It stays silent for Gemini. The list of known keys is built partly from the node's own field names, and this node declares gemini.outputTokens, so outputTokens counts as a real key.
The repo ships a pipeline with this key. pipelines/git_agent_example.pipe sets "outputTokens": 8192 on its custom Gemini node, and that 8,192 is ignored today. #2152 mentions an older pipeline that "carried outputTokens in every llm_gemini node"; the field itself is the likely source.
Steps to Reproduce
- Load
pipelines/git_agent_example.pipe, or add a Gemini node, pick the Custom profile, set a model, and set Output Tokens to 8192.
- Run the pipeline with debug logging on. ChatBase logs
Output tokens : 65536, not 8192.
- Ask for a long answer. The request to Google carries no
max_output_tokens, so the answer can run past 8,192 tokens.
Without a running engine, resolving the example's custom block through the real Config.getNodeConfig (with getServiceDefinition stubbed to return the node's services.json) gives:
{'model': 'gemini-3.1-flash-lite-preview', 'modelTotalTokens': 1000000, 'modelOutputTokens': 65536, 'outputTokens': 8192}
No warning is logged.
Expected vs Actual Behavior
A value typed into Output Tokens should become the node's output limit. ChatBase should use it, and the request to Google should carry it as max_output_tokens.
Today the value sits in the config as outputTokens, ChatBase uses 65,536, and Google receives no limit at all.
Severity
P3 (low, minor inconvenience). Answers are not cut short. The cost is that nobody can cap Gemini output from the node, and the preprocessor sizes chunks from a number the user did not choose.
Platform
All. The code is Python and not platform specific.
Version / Environment
develop at da5ac85.
What a fix needs to do
- Rename the field so it saves as
modelOutputTokens. Use gemini.modelOutputTokens, or the shared modelOutputTokens field id that llm_baidu_qianfan uses.
- Send the limit to Google in
Chat._chat, as max_output_tokens in the config argument of generate_content.
- Decide what happens to pipelines that already saved
outputTokens, and update the example pipeline in the same change.
Pipelines that saved the old key
Renaming loses no working value, because the old value never worked. What changes is what people see.
Without an alias, a saved outputTokens stays ignored and the form shows an empty Output Tokens field, so the number the user typed disappears from view. After the rename, outputTokens is no longer a declared field, so the misnamed-key warning starts to fire for these pipelines. That warning says the key "is ignored", which would be accurate.
With an alias (read outputTokens when modelOutputTokens is absent), old values start to take effect. That is a behavior change. ChatBase refuses to start with a value under 1,024, and any value starts capping output once the driver sends it. The warning would then be wrong, because the key is no longer ignored.
Either choice is defensible, as long as it is made on purpose and tested.
How to test it
Resolve the Custom profile with a modelOutputTokens value and check that getOutputTokens() returns it. Then call the driver with a fake google.genai client and check that generate_content received the limit. Details below.
Related
For your agents
Dense context for an AI agent picking this up. Line numbers are from develop at da5ac85; re-check them before editing. Items marked [verify] were inferred from reading code, not observed.
Where the value goes
Why no limit reaches Google
Chat._chat at nodes/src/nodes/llm_gemini/gemini.py:194 calls self._client.models.generate_content(model=self._model, contents=prompt) with no config.
- The node sets no
_llm and no _native_stream_provider, so ChatBase.chat_string always ends in _chat_with_retries, which calls _chat. There is no other request path to patch.
- The save-time probe in
IGlobal.validateConfig also calls generate_content without a config. Leave it alone; it sends "Hi".
Why the warning stays silent
Config._knownConfigKeys at packages/ai/src/ai/common/config.py:90 adds every field name with its prefix stripped, so gemini.outputTokens contributes outputTokens.
Config._suggestKey at config.py:109 returns None at once when the key is known (lines 127 to 128). Its docstring names this exact dropped-prefix case.
- Confirmed by running the real
config.py against the node's services.json: 'outputTokens' in knownKeys is True and _suggestKey('outputTokens', knownKeys) is None.
Files that carry the old name
What a fix must do
- Rename the field id.
gemini.modelOutputTokens keeps the node's naming; the bare modelOutputTokens matches llm_baidu_qianfan (services.json:135). Update the Custom form's properties list too.
- Pass the limit in
_chat: config=types.GenerateContentConfig(max_output_tokens=self._modelOutputTokens) from google.genai.types, or the equivalent dict. Keep _report_gemini_usage and the GeminiNoTextError handling as they are.
- Choose alias or no alias for
outputTokens (see the human section). If you alias, do it in gemini.py after super().__init__, run the value through validate_max_tokens, and keep the 1,024 floor in mind.
- Move the example pipeline to
modelOutputTokens.
Traps
- Sending a limit changes every catalogue profile, not only Custom. Each will start sending its
modelOutputTokens (65,536, 65,535, 58,982, 32,768 or 8,192 in today's file). A value above what Google accepts for that model fails the call [verify per model]. The values come from the model sync (// openrouter and // sync_models.config.json comments), so a wrong one belongs in tools/sync_models/src/sync_models.config.json, not in a hand edit.
- On Gemini 2.5 and later, thinking tokens count against
max_output_tokens [verify]. A low user value can then end the response before any text, which surfaces as GeminiNoTextError with finish_reason=MAX_TOKENS. That error is already non-retryable, which is correct, but the message should stay readable.
validate_max_tokens clamps the output limit to modelTotalTokens. A custom model with a small context window gets a small output limit, which is right, but test it.
- After the rename,
Config._warnMisnamedKeys will start warning on pipelines that still say outputTokens. If you add an alias, either suppress that warning for this key on this node or reword it, because "is ignored" would be false.
- The stub in
test_gemini_token_metrics.py provides only google.genai.Client. If gemini.py starts importing google.genai.types, that stub needs a types attribute too, or pass a plain dict as config.
nodes/test/mocks/google/genai/__init__.py:33 already accepts config=None in generate_content, so the node's mock test keeps passing whether or not a config is sent. It will not catch a regression unless a test asserts on the argument.
How to test
- Add a unit test in
nodes/test/llm_gemini/. The existing test_gemini_token_metrics.py shows how to load the node module with a fake genai. Build Chat with {'profile': 'custom', 'custom': {'model': 'm', 'apikey': 'k', 'modelOutputTokens': 8192}} and assert getOutputTokens() == 8192.
- Use the same setup with a fake client that records its arguments. Call
_chat('hi') and assert config.max_output_tokens == 8192.
- Check that a catalogue profile sends its catalogue value.
- If you alias, test that
outputTokens: 8192 alone gives 8192, modelOutputTokens wins when both are set, and outputTokens: 512 fails with ChatBase's 1,024 message.
- Resolve
pipelines/git_agent_example.pipe after the change and check the 8,192 reaches ChatBase.
Acceptance criteria
- The Custom form's Output Tokens value is what ChatBase reports and what Google receives as
max_output_tokens.
outputTokens no longer appears in services.json, the README schema table, or the example pipeline.
- The chosen behavior for old
outputTokens values is tested and stated in the PR description.
Out of scope
llm_vision_gemini and accessibility_describe, which build their own generate_content calls.
- Custom forms on other nodes that have no output field at all (Anthropic, OpenAI, xAI, Bedrock, Ollama and others). That gap is real but separate.
Description
The Gemini node's Custom profile has an Output Tokens field. Whatever you type into it has no effect. Two separate problems cause this.
First, the field saves its value under the wrong key. Its id is
gemini.outputTokens, and the engine names a field after the last part of its id, so the value lands in the config asoutputTokens. ChatBase, the base class every LLM node shares, readsmodelOutputTokens. Every other LLM node uses that name. ChatBase therefore ignores the value and uses the 65,536 that the Custom profile carries inservices.json.Second, the Gemini driver never sends an output limit to Google. It calls
generate_contentwith only the model name and the prompt. So even a correctly named field would only change ChatBase's own numbers: the debug log, thegetOutputLengthanswer, and the chunk size the LLM preprocessor derives from it. Google would still apply its own default.The config code already has a warning for this exact mistake.
Config._suggestKeyanswersoutputTokenswith "did you meanmodelOutputTokens?" (#2152, fixed in #2166). It stays silent for Gemini. The list of known keys is built partly from the node's own field names, and this node declaresgemini.outputTokens, sooutputTokenscounts as a real key.The repo ships a pipeline with this key.
pipelines/git_agent_example.pipesets"outputTokens": 8192on its custom Gemini node, and that 8,192 is ignored today. #2152 mentions an older pipeline that "carriedoutputTokensin everyllm_gemininode"; the field itself is the likely source.Steps to Reproduce
pipelines/git_agent_example.pipe, or add a Gemini node, pick the Custom profile, set a model, and set Output Tokens to 8192.Output tokens : 65536, not 8192.max_output_tokens, so the answer can run past 8,192 tokens.Without a running engine, resolving the example's custom block through the real
Config.getNodeConfig(withgetServiceDefinitionstubbed to return the node'sservices.json) gives:No warning is logged.
Expected vs Actual Behavior
A value typed into Output Tokens should become the node's output limit. ChatBase should use it, and the request to Google should carry it as
max_output_tokens.Today the value sits in the config as
outputTokens, ChatBase uses 65,536, and Google receives no limit at all.Severity
P3 (low, minor inconvenience). Answers are not cut short. The cost is that nobody can cap Gemini output from the node, and the preprocessor sizes chunks from a number the user did not choose.
Platform
All. The code is Python and not platform specific.
Version / Environment
developat da5ac85.What a fix needs to do
modelOutputTokens. Usegemini.modelOutputTokens, or the sharedmodelOutputTokensfield id thatllm_baidu_qianfanuses.Chat._chat, asmax_output_tokensin theconfigargument ofgenerate_content.outputTokens, and update the example pipeline in the same change.Pipelines that saved the old key
Renaming loses no working value, because the old value never worked. What changes is what people see.
Without an alias, a saved
outputTokensstays ignored and the form shows an empty Output Tokens field, so the number the user typed disappears from view. After the rename,outputTokensis no longer a declared field, so the misnamed-key warning starts to fire for these pipelines. That warning says the key "is ignored", which would be accurate.With an alias (read
outputTokenswhenmodelOutputTokensis absent), old values start to take effect. That is a behavior change. ChatBase refuses to start with a value under 1,024, and any value starts capping output once the driver sends it. The warning would then be wrong, because the key is no longer ignored.Either choice is defensible, as long as it is made on purpose and tested.
How to test it
Resolve the Custom profile with a
modelOutputTokensvalue and check thatgetOutputTokens()returns it. Then call the driver with a fakegoogle.genaiclient and check thatgenerate_contentreceived the limit. Details below.Related
For your agents
Dense context for an AI agent picking this up. Line numbers are from
developat da5ac85; re-check them before editing. Items marked [verify] were inferred from reading code, not observed.Where the value goes
nodes/src/nodes/llm_gemini/services.json:450and listed on the Custom form atservices.json:555. Its siblingsgemini.modelandgemini.modelTotalTokenswork because their last segment matches a real key.IServices::getFieldNameatpackages/server/engine-lib/engLib/store/services/services.cpp:376, applied to references atservices.cpp:615. Sogemini.outputTokensis saved asoutputTokensin thecustomblock.modelOutputTokens: 65536atservices.json:79.Config.getNodeConfigmerges the user's block over it, so the resolved config has both keys.llm_gemini/IGlobal.pypasses the rawconnConfigtoChat(line 193), so this node resolves once. It is not affected by the double resolution in the companion issue about custom profiles.modelOutputTokensatpackages/ai/src/ai/common/chat.py:106, clamps it at line 124, and raises below 1,024 at lines 127 to 128.getOutputLengthatpackages/ai/src/ai/common/llm_base.py:119, whichpreprocessor_llmcalls atIInstance.py:447and uses as the chunk ceiling at line 197.Why no limit reaches Google
Chat._chatatnodes/src/nodes/llm_gemini/gemini.py:194callsself._client.models.generate_content(model=self._model, contents=prompt)with noconfig._llmand no_native_stream_provider, soChatBase.chat_stringalways ends in_chat_with_retries, which calls_chat. There is no other request path to patch.IGlobal.validateConfigalso callsgenerate_contentwithout a config. Leave it alone; it sends "Hi".Why the warning stays silent
Config._knownConfigKeysatpackages/ai/src/ai/common/config.py:90adds every field name with its prefix stripped, sogemini.outputTokenscontributesoutputTokens.Config._suggestKeyatconfig.py:109returnsNoneat once when the key is known (lines 127 to 128). Its docstring names this exact dropped-prefix case.config.pyagainst the node'sservices.json:'outputTokens' in knownKeysisTrueand_suggestKey('outputTokens', knownKeys)isNone.Files that carry the old name
pipelines/git_agent_example.pipe:66:"outputTokens": 8192in thecustomblock.nodes/src/nodes/llm_gemini/README.md:143: the generated schema table. Regenerate it withnodes:docs-generate; do not edit it by hand.outputTokensoutside tests and the config docstring.What a fix must do
gemini.modelOutputTokenskeeps the node's naming; the baremodelOutputTokensmatchesllm_baidu_qianfan(services.json:135). Update the Custom form'spropertieslist too._chat:config=types.GenerateContentConfig(max_output_tokens=self._modelOutputTokens)fromgoogle.genai.types, or the equivalent dict. Keep_report_gemini_usageand theGeminiNoTextErrorhandling as they are.outputTokens(see the human section). If you alias, do it ingemini.pyaftersuper().__init__, run the value throughvalidate_max_tokens, and keep the 1,024 floor in mind.modelOutputTokens.Traps
modelOutputTokens(65,536, 65,535, 58,982, 32,768 or 8,192 in today's file). A value above what Google accepts for that model fails the call [verify per model]. The values come from the model sync (// openrouterand// sync_models.config.jsoncomments), so a wrong one belongs intools/sync_models/src/sync_models.config.json, not in a hand edit.max_output_tokens[verify]. A low user value can then end the response before any text, which surfaces asGeminiNoTextErrorwithfinish_reason=MAX_TOKENS. That error is already non-retryable, which is correct, but the message should stay readable.validate_max_tokensclamps the output limit tomodelTotalTokens. A custom model with a small context window gets a small output limit, which is right, but test it.Config._warnMisnamedKeyswill start warning on pipelines that still sayoutputTokens. If you add an alias, either suppress that warning for this key on this node or reword it, because "is ignored" would be false.test_gemini_token_metrics.pyprovides onlygoogle.genai.Client. Ifgemini.pystarts importinggoogle.genai.types, that stub needs atypesattribute too, or pass a plain dict asconfig.nodes/test/mocks/google/genai/__init__.py:33already acceptsconfig=Noneingenerate_content, so the node's mock test keeps passing whether or not a config is sent. It will not catch a regression unless a test asserts on the argument.How to test
nodes/test/llm_gemini/. The existingtest_gemini_token_metrics.pyshows how to load the node module with a fakegenai. BuildChatwith{'profile': 'custom', 'custom': {'model': 'm', 'apikey': 'k', 'modelOutputTokens': 8192}}and assertgetOutputTokens() == 8192._chat('hi')and assertconfig.max_output_tokens == 8192.outputTokens: 8192alone gives 8192,modelOutputTokenswins when both are set, andoutputTokens: 512fails with ChatBase's 1,024 message.pipelines/git_agent_example.pipeafter the change and check the 8,192 reaches ChatBase.Acceptance criteria
max_output_tokens.outputTokensno longer appears inservices.json, the README schema table, or the example pipeline.outputTokensvalues is tested and stated in the PR description.Out of scope
llm_vision_geminiandaccessibility_describe, which build their owngenerate_contentcalls.