Skip to content

[Feature Request] Modernize Thai Language Analysis in modules/analysis-common #23151

Description

@kamthorn

Is your feature request related to a problem? Please describe

Thai text analysis in OpenSearch Core (modules/analysis-common) currently provides only a basic ThaiTokenizerFactory (which simply wraps new ThaiTokenizer()) and a default ThaiAnalyzerProvider.

This has several major limitations for production search applications:

  1. Lack of User Dictionary Support: Domain-specific terms (medical, legal, technical jargon, brand names) cannot be injected into the thai tokenizer, resulting in over-segmentation.
  2. Buffer Truncation at 1,024 Characters: As tracked in Apache Lucene issue [BUG] flaky org.opensearch.search.SearchWeightedRoutingIT.testMultiGetWithNetworkDisruption_FailOpenEnabled test #10153 (LUCENE-9112), long Thai paragraphs are split across arbitrary 1,024-character boundaries, corrupting words that span boundaries.
  3. No Character Normalization: Thai orthography allows out-of-order typing (e.g. consonants, tone marks, and vowels in wrong sequence like ก + ่ + า vs ก + า + ่, or decomposed Nikhahit + Sara Aa instead of Sara Am \u0E33), resulting in 0-hit query mismatches.
  4. No Reduplication Support (Maiyamok ๆ): Words with repetition mark ๆ (e.g., เร็วๆ -> เร็ว เร็ว, ดีๆ -> ดี ดี) are ubiquitous in Thai, but standard tokenizers leave ๆ as an unindexed symbol or strip it.

Describe the solution you'd like

Recently, 5 PRs addressing all of these core issues have been contributed and merged into upstream Apache Lucene core:

We would like to propose enhancing OpenSearch's modules/analysis-common:

  1. thai_char_filter / thai_normalizer: Expose Thai character normalization as both a char filter and a normalizer.
  2. thai_repeat: Expose the token filter for Maiyamok ๆ reduplication.
  3. User Dictionary & Buffer Settings: Allow user_dictionary, user_dictionary_rules, and buffer_size in thai tokenizer settings.
  4. Modern thai Analyzer: Incorporate normalization and repeat filtering into the pre-built thai analyzer.

Related component

Search:Relevance

Describe alternatives you've considered

Using custom third-party plugins (such as opensearch-analysis-thaibreak), which provides custom Viterbi segmentation and token filters, but basic out-of-the-box Thai text analysis should work reliably in OpenSearch Core without requiring external plugins for standard normalization and dictionary support.

Additional context

I have authored the 5 upstream PRs in Apache Lucene and am happy to contribute and submit the pull request for this feature to OpenSearch Core.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions