NLP Processing Stages Explained: Lexical to Pragmatic Analysis with a Worked Example

Follow one two-sentence input through tokenisation, morphology, parsing, meaning, coreference and contextual interpretation, with every intermediate result visible.

KnowledgeGate Team

Exam prep & CS education

Updated 18 Sep 20265 min read

Questions may ask for a token, lemma, parse, sense or referent without naming the NLP stage. The running example starts with Riya depositing $5,000 in the bank. She later checked the balance.

Related reading: NLP MCQs and lexical analysis.

1. What the NLP processing stages are and why their boundaries matter

NLP computationally analyses or generates human language. NLP processing can proceed from input and sentence segmentation to lexical and morphological processing, syntactic analysis, semantic analysis, discourse integration and pragmatic interpretation.

Stage

Input

Output

Running-example checkpoint

Segmentation

document

sentences

two sentences

Lexical and morphological

sentences

tokens, forms, lemmas

13 tokens including two full stops; deposit, check

Syntactic

tagged tokens

grammatical relations

nsubj(deposited, Riya)

Semantic

sentence structure

sentence meaning

financial-institution sense of bank

Discourse

sentence meanings

cross-sentence links

{Riya, She}

Pragmatic

discourse plus situation

intended interpretation

account sense of the balance

The stages ask: What are the units? How are they formed? How are they related? What do they mean? How do sentences connect? What is intended?

Tokenisation, POS tagging, NER, parsing, sense disambiguation and coreference are tasks. Lexical, syntactic, semantic and discourse are levels. After reproducing the trace, the GATE CS Exam Preparation Courses & Test Series links this topic to a broader core CS revision path.

2. Lexical and morphological processing: sentences, tokens, forms and lemmas

For this trace, terminal punctuation is separate; $ plus comma-grouped digits stays together, and the symbol denotes USD. S1 is [Riya, deposited, $5,000, in, the, bank, .], 7 tokens; S2 is [She, later, checked, the, balance, .], 6. Thus 7 + 6 = 13 including punctuation and 13 - 2 = 11 without full stops. Tokenizers may split $5,000.

Surface form

Normalised form

Lemma and features

Riya

riya

lemma Riya, PROPN

deposited

deposited

lemma deposit, VERB, past

$5,000

$5,000

NUM/MONEY, value 5000 USD

She

she

lemma she, PRON

checked

checked

lemma check, VERB, past

Lowercasing is normalisation; lemmatisation uses vocabulary and morphology. Suffix stripping can map deposited to deposit and checked to check, but crudely maps studies to studie; a lemmatiser returns study. Preserve casing and offsets to identify Riya. Lexical Analysis in Compiler Design shares lexeme and token terminology. Natural language lacks a formal token grammar and adds ambiguity, scripts and domain conventions.

3. Syntactic analysis: POS tags, phrases and dependency structure

The stated tagging convention gives:

Riya/PROPN deposited/VERB $5,000/NUM in/ADP the/DET bank/NOUN ./PUNCT

She/PRON later/ADV checked/VERB the/DET balance/NOUN ./PUNCT

POS categories depend on context: bank is a noun here, but a verb in Pilots bank the aircraft. Dependencies are:

  • S1: root(deposited), nsubj(deposited, Riya), obj(deposited, $5,000), obl(deposited, bank), case(bank, in), det(bank, the).

  • S2: root(checked), nsubj(checked, She), advmod(checked, later), obj(checked, balance), det(balance, the).

These labels follow the stated convention; treebanks may differ. In the sentence Riya saw the man with a telescope. VP attachment means Riya used it; NP attachment means the man had it. Given score(A)=0.62 and score(B)=0.38, select A since 0.62 > 0.38; scores do not prove intent. Parsing in Compiler Design: Top-Down and Bottom-Up Explained supplies background. Natural language may retain several parses.

Two-lane diagram showing token counts for both sentences and their UPOS tags with dependency arcs.

4. Worked pipeline: carry one input from lexical form to meaning

Step

Preserved output

Raw

Riya deposited $5,000 in the bank. She later checked the balance.

Segmented

S1 and S2 as shown above

Tokenised

7 tokens plus 6 tokens

Normalised

riya, 5000 USD, she; original text retained

Morphology

deposited -> deposit; checked -> check, both past

POS/NER

Riya/PROPN = PERSON; $5,000/NUM = MONEY(value=5000, currency=USD)

Syntax

roots deposited, checked; dependencies above

Sentence semantics

deposit(e1) and check(e2) event structures

Document context

{Riya, She}, account balance, After(e2, e1)

POS gives grammatical category; NER gives entity class. Meanings are deposit(e1), Agent(e1, Riya), Theme(e1, 5000_USD), Destination(e1, bank_financial); and check(e2), Agent(e2, Riya), Theme(e2, account_balance), After(e2, e1). Discourse supplies She -> Riya, later supplies order, context selects account_balance.

The trace preserves every textual or declared value: the amount stays 5000 USD, roots stay deposited/checked, and Riya stays one entity. In Riya sat on the bank and watched the river., bank remains NOUN but becomes river_edge. POS cannot settle sense.

5. Semantic, discourse and pragmatic analysis: three kinds of context

Semantic analysis derives entities, roles, predicates and senses. Use priors P(financial)=P(river_edge)=0.5 and likelihood pairs (deposited,balance)=(0.8,0.7) for financial, (0.05,0.1) for river edge, assuming conditional independence:

  • Financial: 0.5 x 0.8 x 0.7 = 0.28.

  • River edge: 0.5 x 0.05 x 0.1 = 0.0025.

  • Denominator: 0.28 + 0.0025 = 0.2825.

  • Posteriors: 0.28 / 0.2825 = 0.9912 and 0.0025 / 0.2825 = 0.0088, rounded to four decimals.

Financial wins; these probabilities are illustrative, not universal.

Discourse links utterances: {Riya, She} is coreference and later gives After(e2, e1). the balance instead bridges to the account introduced by the deposit.

Pragmatics adds intention, situation and shared knowledge. Can you check the balance? requests action, not an ability test. Beside a scale, the balance may mean the instrument. Context need not guarantee one reading.

Semantic diagram: deposit and check events, bank sense scores and posteriors, She to Riya coreference, and balance bridged to an account.

6. How exams test NLP processing stages

Questions can ask you to order levels, count tokens, distinguish stemming and lemmatisation, label POS or NER, choose attachments, map tasks, calculate scores, or resolve reference:

  1. The declared tokenizer gives 13 tokens including punctuation and 11 without it.

  2. deposited -> deposit is lemmatisation when grammatical analysis supplies the dictionary form; lowercasing is normalisation.

  3. bank/NOUN is syntax; bank_financial is a semantic sense.

  4. Riya = PERSON is NER; Riya/PROPN is POS tagging.

  5. Attaching with a telescope to the VP gives the instrument reading; attaching it to the NP gives the man-has-telescope reading.

  6. She -> Riya is cross-sentence coreference; connecting the balance to an account is bridging plus contextual interpretation.

Likelihood products 0.56 and 0.005 preserve the ranking under equal priors. The posterior is 0.28 / 0.2825 = 0.9912. For only the winner, unnormalised scores suffice because the denominator is shared.

7. Common NLP-stage traps and their precise corrections

  • Token counts: state the tokeniser, count output.

  • Normalisation versus morphology: lowercasing changes presentation; checked -> check derives a lemma. Keep offsets.

  • Stemming versus lemmatisation: rules may emit non-words; lemmatisers use vocabulary and grammar.

  • POS versus NER: Riya/PROPN and Riya/PERSON coexist.

  • Syntax versus semantics: bank remains a noun across both senses.

  • Three contexts: sense is semantic, She -> Riya is discourse, a request reading is pragmatic.

  • Independent boxes: models may annotate jointly; levels classify information.

8. NLP processing stages: the short version and next step

Lexical processing finds units and forms; syntax builds relations; semantics builds meaning; discourse links sentences; pragmatics selects context. Checkpoints: 13 tokens, lemmas deposit/check, roots deposited/checked, posterior 0.9912, {Riya, She}, After(e2, e1).

Replace S2 with She later walked along the river. The pronoun still resolves to Riya and later orders events. River evidence conflicts with the financial reading but does not rewrite the deposit. With new likelihoods, recompute from the supplied values instead of inventing probabilities.

Continue with Artificial Intelligence (AI) for broader practical AI, or GATE Guidance by Sanchit Sir to organise core CS. Identify the required representation before choosing a method.