Skip to content

Support user dictionary in ThaiTokenizer and ThaiTokenizerFactory (#16721) - #16722

Merged
rmuir merged 5 commits into
apache:mainfrom
kamthorn:feature/thai-custom-dictionary
Sep 28, 2026
Merged

rmuir merged 5 commits into
apache:mainfrom
kamthorn:feature/thai-custom-dictionary

Conversation

@kamthorn

Copy link
Copy Markdown
Contributor

Description

Closes #16721

This PR introduces user/custom dictionary support to ThaiTokenizer and ThaiTokenizerFactory, allowing users to supply domain-specific terms, proper nouns, jargon, and loanwords (e.g. "พารากอน", "คลาวด์เนทีฟ", "คนขับรถ") that take precedence over default BreakIterator segmentation boundaries.

Motivation:

ThaiTokenizer uses JRE's BreakIterator with a fixed built-in dictionary. In domain-specific search use cases (e.g. e-commerce, technical documentation, enterprise search), words not present in the default dictionary or multi-word compound terms are often fragmented into sub-words or single characters:

  • "ไปพารากอนกัน" was fragmented into ["ไป", "พา", "รา", "กอน", "กัน"]- "คนขับรถ"was fragmented into["คน", "ขับ", "รถ"]
    Providing user dictionary support brings ThaiTokenizer in line with other language tokenizers in Lucene (such as Kuromoji for Japanese and Nori for Korean).

Changes:

  1. ThaiTokenizer:
    • Added constructors accepting CharArraySet userDictionary.
    • In incrementWord(), when a user dictionary is provided, prefix matches from userDictionary take precedence, cleanly spanning custom terms and adjusting segmentation boundaries.
  2. ThaiTokenizerFactory:
    • Implemented ResourceLoaderAware.
    • Added support for the dictionary parameter (e.g. <tokenizer class="solr.ThaiTokenizerFactory" dictionary="custom_words.txt"/>).
  3. Tests:
    • Added TestThaiTokenizer testing default segmentation, user dictionary term recognition, multi-term offset correctness, and empty dictionary handling.
    • Updated TestThaiTokenizerFactory with testCustomDictionary() using a test dictionary resource (customThaiDictionary.txt).
  4. Backwards Compatibility:
    • Retained all existing constructors and behaviors when no user dictionary is supplied.

@kamthorn
kamthorn force-pushed the feature/thai-custom-dictionary branch from 06d633c to 6b0b916 Compare September 26, 2026 10:44
Comment thread lucene/analysis/common/src/java/org/apache/lucene/analysis/th/ThaiTokenizer.java Outdated
@rmuir
rmuir merged commit 6b44b4c into apache:main Sep 28, 2026
13 checks passed
@kamthorn

Copy link
Copy Markdown
Contributor Author

Hi @rmuir and everyone — thank you for merging the Thai analysis improvements (#16717, #16718, #16720, #16722, #16727). Much appreciated.

I would like to ask about a possible backport to branch_10x.

These changes currently only exist on main (11.0.0, which requires Java 25). branch_10x still contains just the original 4 Thai files (ThaiAnalyzer, ThaiTokenizer, ThaiTokenizerFactory, package-info), so downstream consumers that cannot move to Java 25 yet get none of the new Thai functionality.

The concrete case is OpenSearch: it builds against Lucene 10.5.1 and its minimum JDK is still 21 (opensearch-project/OpenSearch#21174, the JDK 25 bump, is open and facing a Bouncy Castle FIPS objection). So OpenSearch cannot consume Lucene 11 yet, and consequently cannot use ThaiCharFilter/ThaiNormalizer, the curated stopwords, ThaiRepeatFilter, the ThaiTokenizer user dictionary, or the configurable buffer size. Thai search quality there is stuck on the 10.x Thai code.

I noticed that new features are backported to branch_10x when appropriate — e.g. #16680 ("Add type to compound word", a 10.6.0 New Feature) was backported on 2026-09-28 (0157353e). So I wanted to ask whether the Thai changes could be considered for a 10.x backport too.

I understand that backporting new features is a judgement call, so to keep the ask narrow and reviewable, the highest-value item is probably #16722 (user_dictionary on ThaiTokenizer/ThaiTokenizerFactory): self-contained, purely additive, no file-format or breaking public-API change. The rest are similarly additive and analysis-only:

None of them change the index format, so they should be BWC-safe on 10.x.

I am happy to prepare the backport PR(s) against branch_10x myself if that helps — just say the word, or let me know if you would rather handle it as committers. And if the answer is simply "new features don't get backported", that is completely fine; I mainly want to know so I can advise OpenSearch users accurately.

Thanks!

@rmuir

rmuir commented Sep 28, 2026

Copy link
Copy Markdown
Member

@kamthorn I'll backport, no worries. easiest for me just to do it with git merge.

rmuir pushed a commit that referenced this pull request Sep 28, 2026
…6721) (#16722)

* Support user dictionary in ThaiTokenizer and ThaiTokenizerFactory (#16721)

* Update CHANGES.txt for #16722

* Apply spotless formatting to ThaiTokenizer and TestThaiTokenizerFactory

* Simplify user dictionary min/max word length calculation per review suggestion
@rmuir

rmuir commented Sep 28, 2026

Copy link
Copy Markdown
Member

@kaivalnp

Copy link
Copy Markdown
Contributor

FYI this might be causing some test failures: #16736

@rmuir

rmuir commented Sep 28, 2026

Copy link
Copy Markdown
Member

@kaivalnp I'll take a look

@rmuir

rmuir commented Sep 29, 2026

Copy link
Copy Markdown
Member

I was hoping it would be something exciting, but was just the empty string of course.

b304eea

@javanna

javanna commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Hey, thanks for these improvements and thanks for merging them @rmuir !

The new components are a great addition, they can be used in custom analyzers and they were not available before.

I was wondering about the users impact of the set of changes made across those 5 PRs that have been backported. I believe besides providing new building blocks, they modify previously existing, built-in components (thai analyzer, thai tokenizer and stopwords). Is my understanding correct? Is this an ok change to make in a minor? I believe upon upgrade users will get different tokens with the same input, doesn't that require reindexing to ensure predictable results?

@rmuir

rmuir commented Oct 2, 2026

Copy link
Copy Markdown
Member

@javanna yes, I think it is ok. For example, this user dictionary functionality is just a new feature.

Users impacted by the other changes, well, if they are impacted by them, then they were getting problematic results before for said documents.

It isn't a database, its full text, so when we look at things like wrong segmentation, it has to be thought of from that perspective.

javanna added a commit to elastic/elasticsearch that referenced this pull request Oct 2, 2026
Expose `thai_normalization` and `thai_repeat` token filters, as well as `thai` char filter.

- apache/lucene#16717
- apache/lucene#16718
- apache/lucene#16720
- apache/lucene#16722
- apache/lucene#16727

This is the minimum set of changes to make to address build failures against the `lucene_snapshot` branch. The Lucene PRs above include also changes that modify existing thai analysis components like `thai` analyzer and `thai` tokenizer that require reindexing. I am opening a followup release blocker issue for those (#160854).
javanna added a commit to elastic/elasticsearch that referenced this pull request Oct 2, 2026
Expose `thai_normalization` and `thai_repeat` token filters, as well as `thai` char filter.

- apache/lucene#16717
- apache/lucene#16718
- apache/lucene#16720
- apache/lucene#16722
- apache/lucene#16727

This is the minimum set of changes to make to address build failures against the `lucene_snapshot` branch. The Lucene PRs above include also changes that modify existing thai analysis components like `thai` analyzer and `thai` tokenizer that require reindexing. I am opening a followup release blocker issue for those (#160854).
@javanna

javanna commented Oct 2, 2026

Copy link
Copy Markdown
Contributor

I understand about the new features, but it seemed to me that the changes have also modified the output of the thai analyzer, for instance by adding a char filter that was not there before. Is that to be considered a bugfix then because the analyzer was previously behaving incorrectly?

@rmuir

rmuir commented Oct 2, 2026

Copy link
Copy Markdown
Member

yeah, i mean if your output has changed, that means it wasn't working for you before, because e.g. words weren't getting segmented properly / won't match.

its not changes on the order of "oh we added a new stemmer and analyzer wasn't stemming before".

The changes are pretty well documented so you can understand what is happening here.

can't view full text analysis from a lens of "oh you changed the output, thats a break!!!". Not a database. We've got to be able to improve and fix bugs.

@javanna

javanna commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Agreed. Not arguing against this, just trying to understand the scope and reasoning. Thanks for explaining.

Perhaps one last question: how have we communicated this to users in the past with similar changes? We could add some wording to the changes entry about the need to reindex? Or is that not necessary? I worry that users may get silent issues otherwise. For majors we have a migrate guide but that's not the case for minors I believe.

@rmuir

rmuir commented Oct 2, 2026

Copy link
Copy Markdown
Member

its not gonna make anything worse. for example, let's take the tokenizer change.

for the tokenizer change, i made it try to "only break on sentences" because most CJK models are trained on whole sentences. So it was just a bias there. But for thai it doesn't make sense: the way the dictionary-based-break-iterator works, it doesn't make use of a whole sentence. So for thai this was just causing truncated analysis (like "words split in half").

Hence, if before you were getting words truncated due to that, its not going to make anything worse to no longer do it. if you want to take advantage of the improved analysis, you have to reindex in order to get the benefits. then you won't have such truncated tokens anymore.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Support user dictionary in ThaiTokenizer and ThaiTokenizerFactory

4 participants