Repository navigation
Add ThaiRepeatFilter to handle Thai Maiyamok (ๆ) reduplication (#16719) - #16720
Merged
Merged
Conversation
…ffset tests - Ensure both the primary term and repeated Maiyamok tokens receive precise, non-overlapping start/end offsets. - For attached Maiyamok (e.g. 'เร็วๆ'), set primary token endOffset to boundary before 'ๆ' and set repeated token offsets to match the corresponding 'ๆ' character(s). - For multiple standalone Maiyamok (e.g. 'ๆๆ'), partition token offsets evenly across each repetition. - Update TestThaiRepeatFilter with exact offset assertions and add randomized testing with checkRandomData.
…ock for JPMS discovery
Member
|
Looks good to me.
I guess 2 instead of 1 could also be seen as distortion when the reduplication is just noun pluralization. But in experiments with Indonesian at least, I saw no benefits for removing it, so the PR makes sense to me (I am less familiar with how it works in Thai) |
rmuir
pushed a commit
that referenced
this pull request
Sep 28, 2026
… (#16720) * Add ThaiRepeatFilter to handle Thai Maiyamok (ๆ) reduplication (#16719) * Fix OffsetAttribute handling in ThaiRepeatFilter and add randomized offset tests - Ensure both the primary term and repeated Maiyamok tokens receive precise, non-overlapping start/end offsets. - For attached Maiyamok (e.g. 'เร็วๆ'), set primary token endOffset to boundary before 'ๆ' and set repeated token offsets to match the corresponding 'ๆ' character(s). - For multiple standalone Maiyamok (e.g. 'ๆๆ'), partition token offsets evenly across each repetition. - Update TestThaiRepeatFilter with exact offset assertions and add randomized testing with checkRandomData. * Update CHANGES.txt for #16720 * Apply spotless formatting to ThaiRepeatFilter and tests * Fix: register ThaiRepeatFilterFactory in module-info.java provides block for JPMS discovery
This was referenced Oct 2, 2026
javanna
added a commit
to elastic/elasticsearch
that referenced
this pull request
Oct 2, 2026
Expose `thai_normalization` and `thai_repeat` token filters, as well as `thai` char filter. - apache/lucene#16717 - apache/lucene#16718 - apache/lucene#16720 - apache/lucene#16722 - apache/lucene#16727 This is the minimum set of changes to make to address build failures against the `lucene_snapshot` branch. The Lucene PRs above include also changes that modify existing thai analysis components like `thai` analyzer and `thai` tokenizer that require reindexing. I am opening a followup release blocker issue for those (#160854).
javanna
added a commit
to elastic/elasticsearch
that referenced
this pull request
Oct 2, 2026
Expose `thai_normalization` and `thai_repeat` token filters, as well as `thai` char filter. - apache/lucene#16717 - apache/lucene#16718 - apache/lucene#16720 - apache/lucene#16722 - apache/lucene#16727 This is the minimum set of changes to make to address build failures against the `lucene_snapshot` branch. The Lucene PRs above include also changes that modify existing thai analysis components like `thai` analyzer and `thai` tokenizer that require reindexing. I am opening a followup release blocker issue for those (#160854).
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
Closes #16719
This PR introduces
ThaiRepeatFilterandThaiRepeatFilterFactorytolucene/analysis/common, providing support for Thai Maiyamok (ๆ/U+0E46), which denotes word repetition (reduplication).Motivation:
In Thai, Maiyamok (
ๆ) is used after a word or phrase to repeat it:"เร็วๆ"/"เร็ว ๆ"->"เร็ว เร็ว"(very fast)"มากๆ"/"มาก ๆ"->"มาก มาก"(very much)"เด็กๆ"->"เด็ก เด็ก"(children)When tokenized by segmenters like
ThaiTokenizer(based onBreakIterator) orICUTokenizer,ๆis emitted as a standalone token (["เร็ว", "ๆ"]). This distorts term frequency (1 instead of 2), breaks phrase queries (e.g."เร็ว เร็ว"does not match["เร็ว", "ๆ"]), and indexesๆas a meaningless token. In other tokenization pipelines (e.g.WhitespaceTokenizer),ๆmay remain attached at the end of the word (["เร็วๆ"]).Features:
ThaiRepeatFilter:["เร็ว", "ๆ"]\ ->["เร็ว", "เร็ว"]`).["มาก", "ๆ", "ๆ"]\ ->["มาก", "มาก", "มาก"]`).["เร็วๆ"]\ ->["เร็ว", "เร็ว"],"มากๆๆ"\ ->["มาก", "มาก", "มาก"]).ThaiRepeatFilterFactory:thaiRepeat. Registered inMETA-INF/services/org.apache.lucene.analysis.TokenFilterFactory.ThaiAnalyzer:ThaiRepeatFilterinto the analysis pipeline.TestThaiRepeatFiltercovering standalone, attached, multiple, sentence, dangling, andThaiTokenizerintegration.TestThaiRepeatFilterFactory.TestThaiAnalyzer.