Skip to content

Allow configurable buffer size in SegmentingTokenizerBase and recognize whitespace as safe boundary in ThaiTokenizer (#10153) - #16727

Merged
rmuir merged 1 commit into
apache:mainfrom
kamthorn:fix/thai-tokenizer-buffer-safe-end
Sep 28, 2026
Merged

rmuir merged 1 commit into
apache:mainfrom
kamthorn:fix/thai-tokenizer-buffer-safe-end

Conversation

@kamthorn

Copy link
Copy Markdown
Contributor

Description

Fixes #10153 (LUCENE-9112).

SegmentingTokenizerBase uses a hardcoded 1,024-character buffer (BUFFERMAX = 1024) and only recognizes newlines (\r, \n, \u0085, \u2028, \u2029) as unambiguous sentence break positions in isSafeEnd().

In Thai, spaces are used as clause/sentence boundaries rather than newlines, and long paragraphs frequently exceed 1,024 characters without a newline. When findSafeEnd() cannot find a newline in the 1,024-character buffer, usableLength is set to length (1024), cutting text abruptly at the buffer boundary. This causes Thai words spanning across the boundary (e.g. มหาวิทยาลัย spanning indices 1020–1031) to be truncated into invalid word fragments (มหา and วิทยาลัย).

Solution

  1. Configurable Buffer Size in SegmentingTokenizerBase:

  2. Thai Safe Boundary Detection in ThaiTokenizer:

    • Overrode isSafeEnd(char ch) in ThaiTokenizer to recognize whitespace (Character.isWhitespace(ch)).
    • In Thai orthography, whitespace is always an unambiguous word/clause boundary (words never contain whitespace).
    • When refilling, the tokenizer safely stops at the last whitespace before buffer exhaustion, preventing words from being split across buffer boundaries.
    • Exposed bufferSize in ThaiTokenizer constructors and ThaiTokenizerFactory (bufferSize argument).
  3. Tests:

    • Added testCustomBufferSize() in TestSegmentingTokenizerBase with a small buffer size (16).
    • Added testLongTextAcrossBufferBoundary() in TestThaiTokenizer verifying that words spanning across the default 1,024-char boundary remain intact.
    • Added testCustomBufferSize() and testFactoryWithBufferSize() in TestThaiTokenizer.

@rmuir
rmuir merged commit dda6bd3 into apache:main Sep 28, 2026
13 checks passed
rmuir pushed a commit that referenced this pull request Sep 28, 2026
…ze whitespace as safe boundary in ThaiTokenizer (#10153) (#16727)
@rmuir rmuir added this to the 10.6.0 milestone Sep 28, 2026
javanna added a commit to elastic/elasticsearch that referenced this pull request Oct 2, 2026
Expose `thai_normalization` and `thai_repeat` token filters, as well as `thai` char filter.

- apache/lucene#16717
- apache/lucene#16718
- apache/lucene#16720
- apache/lucene#16722
- apache/lucene#16727

This is the minimum set of changes to make to address build failures against the `lucene_snapshot` branch. The Lucene PRs above include also changes that modify existing thai analysis components like `thai` analyzer and `thai` tokenizer that require reindexing. I am opening a followup release blocker issue for those (#160854).
javanna added a commit to elastic/elasticsearch that referenced this pull request Oct 2, 2026
Expose `thai_normalization` and `thai_repeat` token filters, as well as `thai` char filter.

- apache/lucene#16717
- apache/lucene#16718
- apache/lucene#16720
- apache/lucene#16722
- apache/lucene#16727

This is the minimum set of changes to make to address build failures against the `lucene_snapshot` branch. The Lucene PRs above include also changes that modify existing thai analysis components like `thai` analyzer and `thai` tokenizer that require reindexing. I am opening a followup release blocker issue for those (#160854).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

SegmentingTokenizerBase splits terms that occupy 1024th positions in text [LUCENE-9112]

2 participants