Skip to main content
Prof. Deptii Chaudhari, Department of Computer Engineering, I2IT
Lecture Notes - Finite State Automata for NLP
Finite state automata (deterministic and nondeterministic finite automata) provide decision
regarding acceptance and rejection of a string while transducers provide some output for a given
input. Thus, the two machines are quite useful in language processing tasks.
Finite state automata are useful in deciding whether a given word belongs to a particular language
or not. Similarly, transducers are useful in parsing and generation of words from their lexical form.
Areas where finite-state methods have been proved to be particularly efficient are
• phonological
• morphological
• syntactic analyses
• tokenization
• shallow parsing
• word sense disambiguation
• spell checking/correcting
Each path in a finite-state network encodes a single string or a set of strings. All paths in such a
network encode a finite-state language or a finite-state relation.
Consider the following figure encoding the language of the type {clear, clever, ear, ever, fat, fatter}.
FSA for Morphological Analysis
Morphology is a branch of linguistics that studies the inner structure of words. It deals with the
ways the words are formed (word formation rules) using the smallest units bearing some lexical or
grammatical meaning. In modeling natural language morphology, it is necessary to distinguish
between surface forms of words and their lemmas. A lemma for a surface form is typically its
dictionary entry form together with some terms that shows the morphological properties of that
form. So, the lemma for the word “bigger” is represented as big+Adj+Comp which shows that the
word is the comparative form of the adjective big.
Prof. Deptii Chaudhari, Department of Computer Engineering, I2IT
There are two processes by which the morphemes can be combined to form words: inflection and
derivation.
Inflection is the process of adding a grammatical affix to a word stem, forming a word of the same
part of speech as stem. Adding plural s to a noun stem is an example in this respect.
Derivation, on the other hand, is the process of adding an affix to a word stem resulting in a word
having a part of speech different from that of the stem. Adding ness to the adjective good, making
a noun as goodness is an example in this respect.
Most of the inflectional endings in English are regular; however, there are some cases in which
there are irregular inflectional morphology at work. Examples in this regard are foot vs. feet, and
take vs. took.
Finite-state automaton for inflectional morphology of English nouns
For some other kinds of complexities in English inflectional morphology, the finite-state technology
can also find rules to represent them. Fox vs. foxes and mouse vs. mice are examples in this respect.
The transducer mapping mice to mouse can be demonstrated as follows:
A finite-state transducer mapping mice to mouse
Prof. Deptii Chaudhari, Department of Computer Engineering, I2IT
An FSA for another fragment of English derivational morphology
This FSA models several derivational facts, such as the well-known generalization that any verb
ending in -ize can be followed by the nominalizing suffix -ation. Thus, since there is a word fossilize,
we can predict the word fossilization by following states q0, q1, and q2. Similarly, adjectives ending
in -al or -able at q5 (equal, formal, realizable) can take the suffix -ity, or sometimes the suffix -ness
to state q6 (naturalness, casualness)
Exercise: Design a Finite State Automata for baa+!
Q: Finite set of states = q0, q1, q2, q3, q4
∑ : Set of input alphabets = {a,b,!}
q0: Start state
F: Set of final states {q4}
(q, i ) defined by the transition table
Prof. Deptii Chaudhari, Department of Computer Engineering, I2IT
a b !
q0 ϕ q1 ϕ
q1 q2 ϕ ϕ
q2 q3 ϕ ϕ
q3 q3 ϕ q4
q4 ϕ ϕ ϕ
Finite State Transducer for NLP
A finite-state transducer is a non-simple finite-state automaton in which at least one arc is labeled
by a symbol pair such as a:b where a designates the symbol on the upper side of the arc and b the
symbol on the lower side.
There are two types of transducers:
• a functional transducer in which there is at most one output for any input, and
• a sequential transducer in which no state has more than one arc having the same symbol on
the input side.
FST as recognizer: Takes a pair of strings and accepts or rejects them
FST as generator: Outputs a pair of strings for a language
FST as translator: Reads a string and outputs another string and Morphological parsing: letters
(input); morphemes (output)
FST as relater: Computes relations between sets
Below Figure demonstrates a lexical transducer through which the relations {<leaf+NN,
leaf>,<leaf+NNS, leaves>, <left+JJ, left>,<leave+NN, leave>,<leave+NNS, leaves>,<leave+VB.
Leave>,<leave+VBZ, leaves>,<leave+VBD, left>}
are encoded.
As figure shows each word has a distinct path. 0 is the epsilon symbol standing for the empty string.
Such lexical transducers are used both for analysis and generation. To analyze the word leaves, for
instance, one
Prof. Deptii Chaudhari, Department of Computer Engineering, I2IT
needs to follow the path containing the symbols l, e, a, v, e and s on the lower side of the arc label.
There are three such paths in the above network as follows:
a) 0 – l – 1 – e – 2 – a – 3 – v – 4 – e – 5 - +NNS:s – 6
b) 0 – l – 1 – e – 2 – a – 3 – v – 4 – e – 5 - +VBZ:s – 6
c) 0 – l – 1 – e – 2 – a – 3 – f:v –8 – +NNS:e – 9 – 0:s – 6