Repository navigation
Allow configurable buffer size in SegmentingTokenizerBase and recognize whitespace as safe boundary in ThaiTokenizer (#10153) - #16727
Merged
Conversation
…ze whitespace as safe boundary in ThaiTokenizer (apache#10153)
This was referenced Oct 2, 2026
javanna
added a commit
to elastic/elasticsearch
that referenced
this pull request
Oct 2, 2026
Expose `thai_normalization` and `thai_repeat` token filters, as well as `thai` char filter. - apache/lucene#16717 - apache/lucene#16718 - apache/lucene#16720 - apache/lucene#16722 - apache/lucene#16727 This is the minimum set of changes to make to address build failures against the `lucene_snapshot` branch. The Lucene PRs above include also changes that modify existing thai analysis components like `thai` analyzer and `thai` tokenizer that require reindexing. I am opening a followup release blocker issue for those (#160854).
javanna
added a commit
to elastic/elasticsearch
that referenced
this pull request
Oct 2, 2026
Expose `thai_normalization` and `thai_repeat` token filters, as well as `thai` char filter. - apache/lucene#16717 - apache/lucene#16718 - apache/lucene#16720 - apache/lucene#16722 - apache/lucene#16727 This is the minimum set of changes to make to address build failures against the `lucene_snapshot` branch. The Lucene PRs above include also changes that modify existing thai analysis components like `thai` analyzer and `thai` tokenizer that require reindexing. I am opening a followup release blocker issue for those (#160854).
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
Fixes #10153 (LUCENE-9112).
SegmentingTokenizerBaseuses a hardcoded 1,024-character buffer (BUFFERMAX = 1024) and only recognizes newlines (\r,\n,\u0085,\u2028,\u2029) as unambiguous sentence break positions inisSafeEnd().In Thai, spaces are used as clause/sentence boundaries rather than newlines, and long paragraphs frequently exceed 1,024 characters without a newline. When
findSafeEnd()cannot find a newline in the 1,024-character buffer,usableLengthis set tolength(1024), cutting text abruptly at the buffer boundary. This causes Thai words spanning across the boundary (e.g.มหาวิทยาลัยspanning indices 1020–1031) to be truncated into invalid word fragments (มหาandวิทยาลัย).Solution
Configurable Buffer Size in
SegmentingTokenizerBase:int bufferSizetoSegmentingTokenizerBase.BUFFERMAX(1024) as default for complete backward compatibility.Thai Safe Boundary Detection in
ThaiTokenizer:isSafeEnd(char ch)inThaiTokenizerto recognize whitespace (Character.isWhitespace(ch)).bufferSizeinThaiTokenizerconstructors andThaiTokenizerFactory(bufferSizeargument).Tests:
testCustomBufferSize()inTestSegmentingTokenizerBasewith a small buffer size (16).testLongTextAcrossBufferBoundary()inTestThaiTokenizerverifying that words spanning across the default 1,024-char boundary remain intact.testCustomBufferSize()andtestFactoryWithBufferSize()inTestThaiTokenizer.