Data Mining Techniques MCQs with Answers and Worked Explanations

Solve 12 data mining MCQs, then check compact explanations of association rules, supervised learning, confusion matrices, clustering, and feature methods.

KnowledgeGate Team

Exam prep & CS education

Updated 17 Sep 20267 min read

Data-mining questions often look like vocabulary tests, but their distractors exploit precise boundaries: classification versus clustering, a mining technique versus a process stage, or recall versus another confusion-matrix cell, especially under time pressure. Work through all 12 questions before opening the explanations, then record the data-mining concept behind every miss. If you need a wider warm-up first, use the broader DBMS MCQ practice, but return here and reason through every option.

Association Rules and Cross-Selling MCQs

Apriori finds frequent itemsets. Association rules turn co-occurrence into statements such as bread -> butter, supporting market-basket cross-selling rather than segmentation.

Question 1

Exam reference: DSSSB, 2018.

Which of the following data mining algorithms is used for association rule mining?

  • A. Decision tree

  • B. Naïve Bayes

  • C. Apriori

  • D. k-means

Answer: C. Apriori. Decision tree and Naïve Bayes classify; k-means clusters.

Take T1={bread,butter}, T2={bread,milk}, T3={bread,butter,milk}, T4={bread,butter}, T5={milk}. {bread,butter} occurs three times, so support is 3/5 = 60%. Bread occurs four times and accompanies butter three times, so confidence for bread -> butter is 3/4 = 75%.

Question 2

Exam reference: UGC NET, Paper 2 (August), 2016.

Discovery of cross sales opportunities is called as _____.

  • A. Association

  • B. Visualization

  • C. Correlation

  • D. Segmentation

Answer: A. Association. The bread -> butter rule suggests a product to cross-sell when bread appears. Segmentation groups similar customers, visualisation presents patterns, and correlation measures statistical co-movement but does not itself generate market-basket rules.

Classification, Clustering, and Process-Boundary MCQs

Idea

What it does

Learning signal

Classification

Predicts a known label

Labelled examples

Clustering

Forms groups

No supplied class labels

Evaluation

Checks a result

Model or pattern output

Question 3

Exam reference: BEL, Probationary Engineer, 2023.

K-Nearest Neighbours algorithm is a _____ classification algorithm that classifies new data point based on the nearest data point.

  • A. Unsupervised

  • B. Supervised

  • C. Reinforcement

  • D. Deep learning

Answer: B. Supervised. KNN needs labelled neighbours. With k=3 and labels {spam, spam, not spam}, the vote predicts spam. Those known labels make it supervised.

Question 4

Exam reference: UGC NET, Paper 2 (November), 2017.

Which of the following is not a Clustering method ?

  • A. K - Mean method

  • B. Self Organizing feature map method

  • C. K - nearest neighbor method

  • D. Agglomerative method

Answer: C. K - nearest neighbor method. K-means, self-organising maps, and agglomerative methods cluster unlabelled observations. KNN transfers a nearby known label. Do not choose by the shared letter K; the setup matters.

Question 5

Exam reference: DSSSB, 2021.

Which of the following is NOT a Data Mining technique?

  • A. Link Analysis

  • B. Database Segmentation

  • C. Predictive Modeling

  • D. Evaluation

Answer: D. Evaluation. The first three options discover structure or make predictions. Evaluation judges whether the result is useful and reliable. Ask: Does it discover or predict? That points to a technique. Does it judge the output? That points to evaluation.

Entropy, Sensitivity, and Confusion-Matrix MCQs

Entropy measures label impurity. Sensitivity uses the actual-positive row, where a false negative is an actual positive predicted as negative.

Question 6

Exam reference: BEL, Probationary Engineer, 2023.

Which of the following statements is/are correct regarding entropy as a measure of uncertainty in data?

I. Low entropy indicates that the data labels are relatively homogeneous.

II. High entropy indicates that the data labels are more mixed or uncertain.

  • A. Both I and II

  • B. Only II

  • C. Neither I nor II

  • D. Only I

Answer: A. Both I and II. Use H=-sum(p log2 p). For {A,A,A,A}, p(A)=1, so H=0 bits. For {A,A,B,B}, both probabilities are 0.5, so H=-(0.5 x -1 + 0.5 x -1)=1 bit. One set is pure; the other mixed.

Question 7

Exam reference: BEL, Probationary Engineer, 2023.

Which of the following formula is used to calculate the sensitivity of a model?

  • A. FP / (TP + FN)

  • B. TN / (TP + FN)

  • C. TP / (TP + FN)

  • D. FN / (TP + FN)

Answer: C. TP / (TP + FN). The model catches TP out of TP+FN actual positives. For TP=80 and FN=20, sensitivity is 80/(80+20) = 80/100 = 0.80 = 80%.

Question 8

Exam reference: BEL, Probationary Engineer, 2023.

Consider the following confusion matrix.

Positive (Predicted)

Negative (Predicted)

Positive (Actual)

100

50

Negative (Actual)

150

9700

What is the value of False Negative in this confusion matrix?

  • A. 9700

  • B. 50

  • C. 100

  • D. 150

Answer: B. 50. Read actual first, predicted second. The actual-positive, predicted-negative cell is FN=50; the others are TP=100, FP=150, and TN=9700. Sensitivity is 100/(100+50) = 2/3, approximately 66.7%.

Confusion matrix for Question 8 showing TP 100, FN 50, FP 150, TN 9700, with sensitivity of about 66.7 percent.

K-Means, COBWEB, and DBSCAN MCQs

K-means uses centroids. COBWEB updates a concept hierarchy incrementally. DBSCAN expands density-connected regions from core points.

Question 9

Exam reference: UGC NET, Paper 2 (June), 2020.

Given below are two statements:

If two variables 𝑉1 and 𝑉2 are used for clustering, then consider the following statements for 𝑘 means clustering with 𝑘=3:

Statement I: If 𝑉1 and 𝑉2 have correlation of 1 the cluster centroid will be in straight line

Statement II: If 𝑉1 and 𝑉2 have correlation of 0 the cluster centroid will be in straight line

In the light of the above statements, choose the correct answer from the options given below

  • A. Both Statement I and Statement II are true

  • B. Both Statement I and Statement II are false

  • C. Statement I is correct but Statement II is false

  • D. Statement I is incorrect but Statement II is true

Answer: C. Statement I is correct but Statement II is false. On V2=2V1, take (1,2),(2,4),(3,6),(7,14),(8,16),(9,18). Group the first three, next two, and last point. Their centroids, (2,4), (7.5,15), and (9,18), all lie on the line.

For zero correlation, use equal-mass clusters centred at (0,2), (-2,-1), and (2,-1). Their mean is (0,0). The V1 x V2 products 0, 2, and -2 cancel, giving zero covariance and correlation, yet the centroids form a triangle.

Two scatter plots for Question 9: collinear k-means centroids on the line V2 = 2V1, and three zero-correlation centroids forming a triangle.

Question 10

Exam reference: UGC NET, Paper 2 (August), 2016.

In Data mining, ______ is a method of incremental conceptual clustering.

  • A. STRING

  • B. COBWEB

  • C. CORBA

  • D. OLAD

Answer: B. COBWEB. Incremental means updating after each instance; conceptual means keeping a hierarchy of attribute-value statistics. COBWEB uses category utility to insert, merge, split, or create a node. CORBA is distributed-object middleware.

Question 11

Exam reference: UGC NET, Paper 2 (December), 2022.

Select the correct order of DBSCAN algorithm.

A. Find recursively all its density connected points and assign them to the same cluster as the core point.

B. Find all the neighbor points with eps and identify the core points with more than MinPts neighbors.

C. Iterate through the remaining unvisited pointed in the dataset.

D. For each core point if it is not already assigned to a cluster, create a new cluster.

Choose the correct answer from the following :

  • A. B, D, C, A

  • B. D, B, C, A

  • C. B, D, A, C

  • D. D, B, A, C

Answer: C. B, D, A, C. Identify neighbours and core points (B), create a cluster for an unassigned core (D), expand through density-connected points (A), then continue with unvisited points (C).

Set eps=1.5 and MinPts=3, counting the point itself. For A(0,0), B(1,0), C(0,1), and D(5,5), distances among A, B, and C are 1, 1, and sqrt(2), all within eps. Each sees all three nearby points, so they form one cluster. D remains noise.

Match Each Task to the Right Technique

Question 12

Exam reference: UGC NET, Paper 2 (June), 2014.

Match the following

List – I

List – II

a. Classification

i. Principal component analysis

b. Clustering

ii. Branch and Bound

c. Feature Extraction

iii. K-nearest neighbour

d. Feature Selection

iv. K-means

Codes :

  • A. a-iii; b-iv; c-ii; d-i

  • B. a-iv; b-iii; c-i; d-ii

  • C. a-iii; b-iv; c-i; d-ii

  • D. a-iv; b-iii; c-ii; d-i

Answer: C. a-iii; b-iv; c-i; d-ii. KNN predicts labels, so it matches classification. K-means forms groups, so it matches clustering. PCA transforms variables into principal components, so it performs feature extraction. Branch and bound searches subsets of existing variables, so it performs feature selection. Extraction creates new features; selection retains a subset of old ones.

Task

Matching technique

Classification

K-nearest neighbour

Clustering

K-means

Feature extraction

Principal component analysis

Feature selection

Branch and bound

Score the Set and Repair the Weak Concept

Use these score bands to choose your next revision step. With 10-12 correct, revise the two explanations that took longest. With 7-9, redo the missed concept group tomorrow. With 0-6, rebuild the classification-versus-clustering and confusion-matrix basics before retaking the set.

Keep an error log with three columns:

Question

Wrong boundary or formula

One-sentence correction

FN cell

Read predicted before actual

Actual positive + predicted negative = FN

Write each correction from memory before checking it.

After repairing the concepts, move to timed practice with the GATE Test Series.

Data Mining Techniques MCQs: The Short Version and Next Step

Keep five anchors clear:

  • Apriori finds association patterns.

  • KNN performs supervised classification.

  • Entropy measures label impurity.

  • Sensitivity is TP/(TP+FN).

  • K-means, COBWEB, and DBSCAN use distinct clustering mechanisms.

If several groups were weak, rebuild the subject map before another attempt. For structured subject learning, continue with GATE Guidance by Sanchit Sir, or browse the wider GATE CS exam preparation catalogue. Then retake the set without looking at the answers and update your error log.