You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
[Feature Request] Modernize Thai Language Analysis in modules/analysis-common #23151
Is your feature request related to a problem? Please describe
Thai text analysis in OpenSearch Core (modules/analysis-common) currently provides only a basic ThaiTokenizerFactory (which simply wraps new ThaiTokenizer()) and a default ThaiAnalyzerProvider.
This has several major limitations for production search applications:
Lack of User Dictionary Support: Domain-specific terms (medical, legal, technical jargon, brand names) cannot be injected into the thai tokenizer, resulting in over-segmentation.
No Character Normalization: Thai orthography allows out-of-order typing (e.g. consonants, tone marks, and vowels in wrong sequence like ก + ่ + า vs ก + า + ่, or decomposed Nikhahit + Sara Aa instead of Sara Am \u0E33), resulting in 0-hit query mismatches.
No Reduplication Support (Maiyamok ๆ): Words with repetition mark ๆ (e.g., เร็วๆ -> เร็ว เร็ว, ดีๆ -> ดี ดี) are ubiquitous in Thai, but standard tokenizers leave ๆ as an unindexed symbol or strip it.
Describe the solution you'd like
Recently, 5 PRs addressing all of these core issues have been contributed and merged into upstream Apache Lucene core:
apache/lucene#16717 - ThaiCharFilter, ThaiNormalizationFilter, and ThaiNormalizer
apache/lucene#16718 - Modern Thai Stopwords (prevents over-filtering of compound words)
apache/lucene#16720 - ThaiRepeatFilter (Maiyamok ๆ reduplication with correct offsets)
We would like to propose enhancing OpenSearch's modules/analysis-common:
thai_char_filter / thai_normalizer: Expose Thai character normalization as both a char filter and a normalizer.
thai_repeat: Expose the token filter for Maiyamok ๆ reduplication.
User Dictionary & Buffer Settings: Allow user_dictionary, user_dictionary_rules, and buffer_size in thai tokenizer settings.
Modern thai Analyzer: Incorporate normalization and repeat filtering into the pre-built thai analyzer.
Related component
Search:Relevance
Describe alternatives you've considered
Using custom third-party plugins (such as opensearch-analysis-thaibreak), which provides custom Viterbi segmentation and token filters, but basic out-of-the-box Thai text analysis should work reliably in OpenSearch Core without requiring external plugins for standard normalization and dictionary support.
Additional context
I have authored the 5 upstream PRs in Apache Lucene and am happy to contribute and submit the pull request for this feature to OpenSearch Core.
Is your feature request related to a problem? Please describe
Thai text analysis in OpenSearch Core (
modules/analysis-common) currently provides only a basicThaiTokenizerFactory(which simply wrapsnew ThaiTokenizer()) and a defaultThaiAnalyzerProvider.This has several major limitations for production search applications:
thaitokenizer, resulting in over-segmentation.ก+่+าvsก+า+่, or decomposed Nikhahit + Sara Aa instead of Sara Am\u0E33), resulting in 0-hit query mismatches.ๆ): Words with repetition markๆ(e.g.,เร็วๆ->เร็ว เร็ว,ดีๆ->ดี ดี) are ubiquitous in Thai, but standard tokenizers leaveๆas an unindexed symbol or strip it.Describe the solution you'd like
Recently, 5 PRs addressing all of these core issues have been contributed and merged into upstream Apache Lucene core:
ThaiCharFilter,ThaiNormalizationFilter, andThaiNormalizerThaiRepeatFilter(Maiyamokๆreduplication with correct offsets)ThaiTokenizerWe would like to propose enhancing OpenSearch's
modules/analysis-common:thai_char_filter/thai_normalizer: Expose Thai character normalization as both a char filter and a normalizer.thai_repeat: Expose the token filter for Maiyamokๆreduplication.user_dictionary,user_dictionary_rules, andbuffer_sizeinthaitokenizer settings.thaiAnalyzer: Incorporate normalization and repeat filtering into the pre-builtthaianalyzer.Related component
Search:Relevance
Describe alternatives you've considered
Using custom third-party plugins (such as
opensearch-analysis-thaibreak), which provides custom Viterbi segmentation and token filters, but basic out-of-the-box Thai text analysis should work reliably in OpenSearch Core without requiring external plugins for standard normalization and dictionary support.Additional context
I have authored the 5 upstream PRs in Apache Lucene and am happy to contribute and submit the pull request for this feature to OpenSearch Core.