NLP Processing Stages Explained: Lexical to Pragmatic Analysis with a Worked Example
Follow one two-sentence input through tokenisation, morphology, parsing, meaning, coreference and contextual interpretation, with every intermediate result visible.
KnowledgeGate Team
Exam prep & CS education

Questions may ask for a token, lemma, parse, sense or referent without naming the NLP stage. The running example starts with Riya depositing $5,000 in the bank. She later checked the balance.
Related reading: NLP MCQs and lexical analysis.
1. What the NLP processing stages are and why their boundaries matter
NLP computationally analyses or generates human language. NLP processing can proceed from input and sentence segmentation to lexical and morphological processing, syntactic analysis, semantic analysis, discourse integration and pragmatic interpretation.
Stage | Input | Output | Running-example checkpoint |
|---|---|---|---|
Segmentation | document | sentences | two sentences |
Lexical and morphological | sentences | tokens, forms, lemmas | 13 tokens including two full stops; |
Syntactic | tagged tokens | grammatical relations |
|
Semantic | sentence structure | sentence meaning | financial-institution sense of |
Discourse | sentence meanings | cross-sentence links |
|
Pragmatic | discourse plus situation | intended interpretation | account sense of |
The stages ask: What are the units? How are they formed? How are they related? What do they mean? How do sentences connect? What is intended?
Tokenisation, POS tagging, NER, parsing, sense disambiguation and coreference are tasks. Lexical, syntactic, semantic and discourse are levels. After reproducing the trace, the GATE CS Exam Preparation Courses & Test Series links this topic to a broader core CS revision path.
2. Lexical and morphological processing: sentences, tokens, forms and lemmas
For this trace, terminal punctuation is separate; $ plus comma-grouped digits stays together, and the symbol denotes USD. S1 is [Riya, deposited, $5,000, in, the, bank, .], 7 tokens; S2 is [She, later, checked, the, balance, .], 6. Thus 7 + 6 = 13 including punctuation and 13 - 2 = 11 without full stops. Tokenizers may split $5,000.
Surface form | Normalised form | Lemma and features |
|---|---|---|
|
| lemma |
|
| lemma |
|
|
|
|
| lemma |
|
| lemma |
Lowercasing is normalisation; lemmatisation uses vocabulary and morphology. Suffix stripping can map deposited to deposit and checked to check, but crudely maps studies to studie; a lemmatiser returns study. Preserve casing and offsets to identify Riya. Lexical Analysis in Compiler Design shares lexeme and token terminology. Natural language lacks a formal token grammar and adds ambiguity, scripts and domain conventions.
3. Syntactic analysis: POS tags, phrases and dependency structure
The stated tagging convention gives:
Riya/PROPN deposited/VERB $5,000/NUM in/ADP the/DET bank/NOUN ./PUNCT
She/PRON later/ADV checked/VERB the/DET balance/NOUN ./PUNCT
POS categories depend on context: bank is a noun here, but a verb in Pilots bank the aircraft. Dependencies are:
S1:
root(deposited),nsubj(deposited, Riya),obj(deposited, $5,000),obl(deposited, bank),case(bank, in),det(bank, the).S2:
root(checked),nsubj(checked, She),advmod(checked, later),obj(checked, balance),det(balance, the).
These labels follow the stated convention; treebanks may differ. In the sentence Riya saw the man with a telescope. VP attachment means Riya used it; NP attachment means the man had it. Given score(A)=0.62 and score(B)=0.38, select A since 0.62 > 0.38; scores do not prove intent. Parsing in Compiler Design: Top-Down and Bottom-Up Explained supplies background. Natural language may retain several parses.

4. Worked pipeline: carry one input from lexical form to meaning
Step | Preserved output |
|---|---|
Raw |
|
Segmented | S1 and S2 as shown above |
Tokenised | 7 tokens plus 6 tokens |
Normalised |
|
Morphology |
|
POS/NER |
|
Syntax | roots |
Sentence semantics |
|
Document context |
|
POS gives grammatical category; NER gives entity class. Meanings are deposit(e1), Agent(e1, Riya), Theme(e1, 5000_USD), Destination(e1, bank_financial); and check(e2), Agent(e2, Riya), Theme(e2, account_balance), After(e2, e1). Discourse supplies She -> Riya, later supplies order, context selects account_balance.
The trace preserves every textual or declared value: the amount stays 5000 USD, roots stay deposited/checked, and Riya stays one entity. In Riya sat on the bank and watched the river., bank remains NOUN but becomes river_edge. POS cannot settle sense.
5. Semantic, discourse and pragmatic analysis: three kinds of context
Semantic analysis derives entities, roles, predicates and senses. Use priors P(financial)=P(river_edge)=0.5 and likelihood pairs (deposited,balance)=(0.8,0.7) for financial, (0.05,0.1) for river edge, assuming conditional independence:
Financial:
0.5 x 0.8 x 0.7 = 0.28.River edge:
0.5 x 0.05 x 0.1 = 0.0025.Denominator:
0.28 + 0.0025 = 0.2825.Posteriors:
0.28 / 0.2825 = 0.9912and0.0025 / 0.2825 = 0.0088, rounded to four decimals.
Financial wins; these probabilities are illustrative, not universal.
Discourse links utterances: {Riya, She} is coreference and later gives After(e2, e1). the balance instead bridges to the account introduced by the deposit.
Pragmatics adds intention, situation and shared knowledge. Can you check the balance? requests action, not an ability test. Beside a scale, the balance may mean the instrument. Context need not guarantee one reading.

6. How exams test NLP processing stages
Questions can ask you to order levels, count tokens, distinguish stemming and lemmatisation, label POS or NER, choose attachments, map tasks, calculate scores, or resolve reference:
The declared tokenizer gives
13tokens including punctuation and11without it.deposited -> depositis lemmatisation when grammatical analysis supplies the dictionary form; lowercasing is normalisation.bank/NOUNis syntax;bank_financialis a semantic sense.Riya = PERSONis NER;Riya/PROPNis POS tagging.Attaching
with a telescopeto the VP gives the instrument reading; attaching it to the NP gives the man-has-telescope reading.She -> Riyais cross-sentence coreference; connectingthe balanceto an account is bridging plus contextual interpretation.
Likelihood products 0.56 and 0.005 preserve the ranking under equal priors. The posterior is 0.28 / 0.2825 = 0.9912. For only the winner, unnormalised scores suffice because the denominator is shared.
7. Common NLP-stage traps and their precise corrections
Token counts: state the tokeniser, count output.
Normalisation versus morphology: lowercasing changes presentation;
checked -> checkderives a lemma. Keep offsets.Stemming versus lemmatisation: rules may emit non-words; lemmatisers use vocabulary and grammar.
POS versus NER:
Riya/PROPNandRiya/PERSONcoexist.Syntax versus semantics:
bankremains a noun across both senses.Three contexts: sense is semantic,
She -> Riyais discourse, a request reading is pragmatic.Independent boxes: models may annotate jointly; levels classify information.
8. NLP processing stages: the short version and next step
Lexical processing finds units and forms; syntax builds relations; semantics builds meaning; discourse links sentences; pragmatics selects context. Checkpoints: 13 tokens, lemmas deposit/check, roots deposited/checked, posterior 0.9912, {Riya, She}, After(e2, e1).
Replace S2 with She later walked along the river. The pronoun still resolves to Riya and later orders events. River evidence conflicts with the financial reading but does not rewrite the deposit. With new likelihoods, recompute from the supplied values instead of inventing probabilities.
Continue with Artificial Intelligence (AI) for broader practical AI, or GATE Guidance by Sanchit Sir to organise core CS. Identify the required representation before choosing a method.
Keep learning

Paging and TLB Explained: Address Translation, EMAT and Exam Traps
Follow one virtual address from its VPN through the TLB to a physical frame, then calculate page-table size, TLB reach and effective memory access time.

Operating System Scenarios: Solve Scheduling, Concurrency, and Page Replacement
Learn one state-trace method for three common OS problem families, then apply it to complete Round Robin, concurrency, FIFO, and LRU examples.

Tower Research Hiring Process: Stage-by-Stage Prep for Quant and Dev Roles
Prepare for a Tower Research application without treating one online account as a universal process. Use this role-led map, worked drills, and seven-day plan.

Capital One Recruitment Process: Stage-by-Stage Guide for India Applicants
Prepare for a Capital One India application with a cautious five-stage map, worked technical and case drills, and a practical 14-hour schedule.