Skip to content

Add ThaiRepeatFilter to handle Thai Maiyamok (ๆ) reduplication (#16719) - #16720

Merged
rmuir merged 6 commits into
apache:mainfrom
kamthorn:feature/thai-repeat-filter
Sep 27, 2026
Merged

rmuir merged 6 commits into
apache:mainfrom
kamthorn:feature/thai-repeat-filter

Conversation

@kamthorn

Copy link
Copy Markdown
Contributor

Description

Closes #16719

This PR introduces ThaiRepeatFilter and ThaiRepeatFilterFactory to lucene/analysis/common, providing support for Thai Maiyamok (ๆ / U+0E46), which denotes word repetition (reduplication).

Motivation:

In Thai, Maiyamok (ๆ) is used after a word or phrase to repeat it:

  • "เร็วๆ" / "เร็ว ๆ" -> "เร็ว เร็ว" (very fast)
  • "มากๆ" / "มาก ๆ" -> "มาก มาก" (very much)
  • "เด็กๆ" -> "เด็ก เด็ก" (children)

When tokenized by segmenters like ThaiTokenizer (based on BreakIterator) or ICUTokenizer, ๆ is emitted as a standalone token (["เร็ว", "ๆ"]). This distorts term frequency (1 instead of 2), breaks phrase queries (e.g. "เร็ว เร็ว" does not match ["เร็ว", "ๆ"]), and indexes ๆ as a meaningless token. In other tokenization pipelines (e.g. WhitespaceTokenizer), ๆ may remain attached at the end of the word (["เร็วๆ"]).

Features:

  • ThaiRepeatFilter:
    • Handles standalone Maiyamok tokens (["เร็ว", "ๆ"]\ -> ["เร็ว", "เร็ว"]`).
    • Handles multiple consecutive Maiyamok (["มาก", "ๆ", "ๆ"]\ -> ["มาก", "มาก", "มาก"]`).
    • Handles attached trailing Maiyamok (["เร็วๆ"]\ -> ["เร็ว", "เร็ว"], "มากๆๆ"\ -> ["มาก", "มาก", "มาก"]).
    • Gracefully discards dangling Maiyamok at stream start with no preceding token.
  • ThaiRepeatFilterFactory:
    • SPI name: thaiRepeat. Registered in META-INF/services/org.apache.lucene.analysis.TokenFilterFactory.
  • ThaiAnalyzer:
    • Integrated ThaiRepeatFilter into the analysis pipeline.
  • Tests:
    • Added TestThaiRepeatFilter covering standalone, attached, multiple, sentence, dangling, and ThaiTokenizer integration.
    • Added TestThaiRepeatFilterFactory.
    • Added test case in TestThaiAnalyzer.

kamthorn and others added 5 commits September 26, 2026 16:19
…ffset tests

- Ensure both the primary term and repeated Maiyamok tokens receive precise, non-overlapping start/end offsets.
- For attached Maiyamok (e.g. 'เร็วๆ'), set primary token endOffset to boundary before 'ๆ' and set repeated token offsets to match the corresponding 'ๆ' character(s).
- For multiple standalone Maiyamok (e.g. 'ๆๆ'), partition token offsets evenly across each repetition.
- Update TestThaiRepeatFilter with exact offset assertions and add randomized testing with checkRandomData.
@rmuir

rmuir commented Sep 27, 2026

Copy link
Copy Markdown
Member

Looks good to me.

This distorts term frequency (1 instead of 2)

I guess 2 instead of 1 could also be seen as distortion when the reduplication is just noun pluralization. But in experiments with Indonesian at least, I saw no benefits for removing it, so the PR makes sense to me (I am less familiar with how it works in Thai)

@rmuir
rmuir merged commit b7c2c20 into apache:main Sep 27, 2026
13 checks passed
rmuir pushed a commit that referenced this pull request Sep 28, 2026
… (#16720)

* Add ThaiRepeatFilter to handle Thai Maiyamok (ๆ) reduplication (#16719)

* Fix OffsetAttribute handling in ThaiRepeatFilter and add randomized offset tests

- Ensure both the primary term and repeated Maiyamok tokens receive precise, non-overlapping start/end offsets.
- For attached Maiyamok (e.g. 'เร็วๆ'), set primary token endOffset to boundary before 'ๆ' and set repeated token offsets to match the corresponding 'ๆ' character(s).
- For multiple standalone Maiyamok (e.g. 'ๆๆ'), partition token offsets evenly across each repetition.
- Update TestThaiRepeatFilter with exact offset assertions and add randomized testing with checkRandomData.

* Update CHANGES.txt for #16720

* Apply spotless formatting to ThaiRepeatFilter and tests

* Fix: register ThaiRepeatFilterFactory in module-info.java provides block for JPMS discovery
@rmuir rmuir added this to the 10.6.0 milestone Sep 28, 2026
javanna added a commit to elastic/elasticsearch that referenced this pull request Oct 2, 2026
Expose `thai_normalization` and `thai_repeat` token filters, as well as `thai` char filter.

- apache/lucene#16717
- apache/lucene#16718
- apache/lucene#16720
- apache/lucene#16722
- apache/lucene#16727

This is the minimum set of changes to make to address build failures against the `lucene_snapshot` branch. The Lucene PRs above include also changes that modify existing thai analysis components like `thai` analyzer and `thai` tokenizer that require reindexing. I am opening a followup release blocker issue for those (#160854).
javanna added a commit to elastic/elasticsearch that referenced this pull request Oct 2, 2026
Expose `thai_normalization` and `thai_repeat` token filters, as well as `thai` char filter.

- apache/lucene#16717
- apache/lucene#16718
- apache/lucene#16720
- apache/lucene#16722
- apache/lucene#16727

This is the minimum set of changes to make to address build failures against the `lucene_snapshot` branch. The Lucene PRs above include also changes that modify existing thai analysis components like `thai` analyzer and `thai` tokenizer that require reindexing. I am opening a followup release blocker issue for those (#160854).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Add ThaiRepeatFilter to handle Thai Maiyamok (ๆ) reduplication

2 participants