Repository navigation
Support user dictionary in ThaiTokenizer and ThaiTokenizerFactory (#16721) - #16722
Conversation
06d633c to
6b0b916
Compare
|
Hi @rmuir and everyone — thank you for merging the Thai analysis improvements (#16717, #16718, #16720, #16722, #16727). Much appreciated. I would like to ask about a possible backport to These changes currently only exist on The concrete case is OpenSearch: it builds against Lucene 10.5.1 and its minimum JDK is still 21 (opensearch-project/OpenSearch#21174, the JDK 25 bump, is open and facing a Bouncy Castle FIPS objection). So OpenSearch cannot consume Lucene 11 yet, and consequently cannot use I noticed that new features are backported to I understand that backporting new features is a judgement call, so to keep the ask narrow and reviewable, the highest-value item is probably #16722 (
None of them change the index format, so they should be BWC-safe on 10.x. I am happy to prepare the backport PR(s) against Thanks! |
|
@kamthorn I'll backport, no worries. easiest for me just to do it with git merge. |
|
FYI this might be causing some test failures: #16736 |
|
@kaivalnp I'll take a look |
|
I was hoping it would be something exciting, but was just the empty string of course. |
|
Hey, thanks for these improvements and thanks for merging them @rmuir ! The new components are a great addition, they can be used in custom analyzers and they were not available before. I was wondering about the users impact of the set of changes made across those 5 PRs that have been backported. I believe besides providing new building blocks, they modify previously existing, built-in components (thai analyzer, thai tokenizer and stopwords). Is my understanding correct? Is this an ok change to make in a minor? I believe upon upgrade users will get different tokens with the same input, doesn't that require reindexing to ensure predictable results? |
|
@javanna yes, I think it is ok. For example, this user dictionary functionality is just a new feature. Users impacted by the other changes, well, if they are impacted by them, then they were getting problematic results before for said documents. It isn't a database, its full text, so when we look at things like wrong segmentation, it has to be thought of from that perspective. |
Expose `thai_normalization` and `thai_repeat` token filters, as well as `thai` char filter. - apache/lucene#16717 - apache/lucene#16718 - apache/lucene#16720 - apache/lucene#16722 - apache/lucene#16727 This is the minimum set of changes to make to address build failures against the `lucene_snapshot` branch. The Lucene PRs above include also changes that modify existing thai analysis components like `thai` analyzer and `thai` tokenizer that require reindexing. I am opening a followup release blocker issue for those (#160854).
Expose `thai_normalization` and `thai_repeat` token filters, as well as `thai` char filter. - apache/lucene#16717 - apache/lucene#16718 - apache/lucene#16720 - apache/lucene#16722 - apache/lucene#16727 This is the minimum set of changes to make to address build failures against the `lucene_snapshot` branch. The Lucene PRs above include also changes that modify existing thai analysis components like `thai` analyzer and `thai` tokenizer that require reindexing. I am opening a followup release blocker issue for those (#160854).
|
I understand about the new features, but it seemed to me that the changes have also modified the output of the |
|
yeah, i mean if your output has changed, that means it wasn't working for you before, because e.g. words weren't getting segmented properly / won't match. its not changes on the order of "oh we added a new stemmer and analyzer wasn't stemming before". The changes are pretty well documented so you can understand what is happening here. can't view full text analysis from a lens of "oh you changed the output, thats a break!!!". Not a database. We've got to be able to improve and fix bugs. |
|
Agreed. Not arguing against this, just trying to understand the scope and reasoning. Thanks for explaining. Perhaps one last question: how have we communicated this to users in the past with similar changes? We could add some wording to the changes entry about the need to reindex? Or is that not necessary? I worry that users may get silent issues otherwise. For majors we have a migrate guide but that's not the case for minors I believe. |
|
its not gonna make anything worse. for example, let's take the tokenizer change. for the tokenizer change, i made it try to "only break on sentences" because most CJK models are trained on whole sentences. So it was just a bias there. But for thai it doesn't make sense: the way the dictionary-based-break-iterator works, it doesn't make use of a whole sentence. So for thai this was just causing truncated analysis (like "words split in half"). Hence, if before you were getting words truncated due to that, its not going to make anything worse to no longer do it. if you want to take advantage of the improved analysis, you have to reindex in order to get the benefits. then you won't have such truncated tokens anymore. |
Description
Closes #16721
This PR introduces user/custom dictionary support to
ThaiTokenizerandThaiTokenizerFactory, allowing users to supply domain-specific terms, proper nouns, jargon, and loanwords (e.g. "พารากอน", "คลาวด์เนทีฟ", "คนขับรถ") that take precedence over defaultBreakIteratorsegmentation boundaries.Motivation:
ThaiTokenizeruses JRE'sBreakIteratorwith a fixed built-in dictionary. In domain-specific search use cases (e.g. e-commerce, technical documentation, enterprise search), words not present in the default dictionary or multi-word compound terms are often fragmented into sub-words or single characters:"ไปพารากอนกัน"was fragmented into["ไป", "พา", "รา", "กอน", "กัน"]-"คนขับรถ"was fragmented into["คน", "ขับ", "รถ"]Providing user dictionary support brings
ThaiTokenizerin line with other language tokenizers in Lucene (such as Kuromoji for Japanese and Nori for Korean).Changes:
ThaiTokenizer:CharArraySet userDictionary.incrementWord(), when a user dictionary is provided, prefix matches fromuserDictionarytake precedence, cleanly spanning custom terms and adjusting segmentation boundaries.ThaiTokenizerFactory:ResourceLoaderAware.dictionaryparameter (e.g.<tokenizer class="solr.ThaiTokenizerFactory" dictionary="custom_words.txt"/>).TestThaiTokenizertesting default segmentation, user dictionary term recognition, multi-term offset correctness, and empty dictionary handling.TestThaiTokenizerFactorywithtestCustomDictionary()using a test dictionary resource (customThaiDictionary.txt).