Intelligent and Rational Agents in AI: PEAS, Worked Decisions and Exam Traps

Learn what makes an AI agent rational, specify its task with PEAS, update a belief using Bayes' rule, and see why coordination matters in multi-agent systems.

KnowledgeGate Team

Exam prep & CS education

Updated 7 Oct 20266 min read

The words agent, intelligence and rationality can sound interchangeable until a question asks whether a failed action can still be rational. The key answer is yes: rationality judges the decision using the percepts and knowledge available at that time, not only the outcome eventually realised. PEAS specifies the task, Bayesian updating changes a route choice after evidence, and a bridge matrix shows why individually rational agents may still need coordination. Use the Skills category as a broader learning route around these foundations.

Intelligent and rational agents: the foundation

An agent receives percepts from an environment through sensors and affects that environment through actuators. A temperature reading at time t is the current percept p_t. The complete history p_1, ..., p_t is the percept sequence, which lets a stateful agent use earlier evidence too.

The abstract agent function is f: P* -> A. It maps every possible percept sequence to an action. Intelligence is the broader capacity to represent, reason, learn, plan, communicate and adapt. Rationality is narrower: choose the action expected to maximise the stated performance measure, given the percept sequence, prior knowledge and available actions.

A sophisticated agent may intelligently optimise the wrong objective. A simple rule may be rational in a simple, fully observable task. Rationality also does not demand formal logic at every step. Logic can represent rules and consequences, probability can represent uncertainty, and search can compare future paths. The Approaches to AI: Four Approaches and a Worked Example compares acting and thinking across human and rational criteria, including the Turing Test. Rational-agent design goes further by specifying PEAS, updating uncertain beliefs and coordinating actions.

PEAS and the four determinants of rationality

PEAS specifies the task environment, not the agent's internal software modules. Consider one courier deciding between a short route S and a covered detour D.

PEAS element

Courier specification

Performance measure

Net decision score: +18 if S is clear and succeeds, -6 if S is blocked, and +8 for D

Environment

Junction J, short route S, covered detour D, and a possible blockage on S

Actuators

Steering and drive controls that execute Take S or Take D

Sensors

Route camera and warning signal W about a possible blockage

Four determinants fix the rational choice here. The performance measure is the score above. The percept history ends in W or no warning. Prior knowledge is P(blocked)=0.20, P(clear)=0.80, P(W|blocked)=0.90, and P(W|clear)=0.20. The available actions are Take S and Take D. Changing any one can change the rational action.

Autonomy does not mean beginning without knowledge. An autonomous agent should increasingly base behaviour on its percepts and experience. If repeated journeys revise the blockage prior or the sensor likelihoods, rational behaviour uses those revised values next time.

Rational-agent worked example: update belief, then choose

Before a warning, calculate the expected score of the short route:

EU(Take S)=0.80(18)+0.20(-6)=14.4-1.2=13.2

The detour has EU(Take D)=8. Since 13.2>8, taking S is rational before the warning.

Now observe W. Bayes' rule first needs the total probability of that warning:

P(W)=0.90(0.20)+0.20(0.80)=0.18+0.16=0.34

Therefore, P(blocked|W)=0.18/0.34=9/17, about 0.5294, while P(clear|W)=0.16/0.34=8/17, about 0.4706. Recompute the action:

EU(Take S|W)=(8/17)(18)+(9/17)(-6)=(144-54)/17=90/17, about 5.29.

Because EU(Take D|W)=8, the warning changes the rational action to Take D. If S later proves clear, the detour remains rational because the decision used the evidence available then. Graph Algorithms: BFS, DFS and Shortest Paths can help generate route alternatives, while beliefs and expected performance decide between uncertain outcomes.

Two-stage courier decision tree: EU(S)=13.2 picks the short route, then after warning W, EU(S|W)=5.29 against EU(D|W)=8 picks the detour.

Task-environment properties that change a rational design

Dimension

Design question

Fully observable / partially observable

Does the percept expose the complete relevant state, or must the agent maintain beliefs about hidden state?

Deterministic / stochastic

Does an action fix its outcome, or must the agent predict a distribution of outcomes?

Episodic / sequential

Is each choice independent, or does it alter later options?

Static / dynamic

Can the environment change while the agent reasons?

Discrete / continuous

Are modelled states and actions countable or continuously varying?

Single-agent / multi-agent

Do other decision makers affect outcomes?

Known / unknown

Are action and observation models known or learned?

The courier snapshot is partially observable because blockage is hidden behind a noisy warning. It is stochastic, sequential because a route fixes later position and options, and discrete because the model has two actions and two blockage states. It is static only for this calculation. If traffic changes during reasoning, the continuing task is dynamic.

These labels change design. Full observability may remove belief state, stochasticity requires probabilities, sequential tasks require future consequences, dynamic settings penalise delay, and unknown models motivate learning. A task is multi-agent only when other decision makers affect it.

Rational agents in a multi-agent environment

Another agent is not merely a moving obstacle. It chooses from its own percepts, knowledge and performance measure, perhaps responding to your expected action. The setting may be cooperative, competitive or mixed.

For two robots approaching a one-lane bridge, let each choose Enter or Yield:

A choice

B choice

Payoff (A,B)

Enter

Enter

(-6,-6)

Enter

Yield

(4,1)

Yield

Enter

(1,4)

Yield

Yield

(0,0)

Without coordination, symmetric robots might both enter for total -12, or both yield for 0. Either asymmetric outcome totals 5. With the public rule “lower robot ID enters first”, lower-ID A enters and B yields, producing (4,1). The convention is rational only for these actions, payoffs and shared knowledge. Individual rationality does not guarantee a rational system result.

Rational-agent traps that change the answer

Trap

Why it fails

Correction

Rational means always successful

Chance can produce a bad result

Judge expected performance at decision time

Rational means omniscient

Future state can be hidden

Use available percepts and knowledge

Rational means human-like

Human imitation is not the criterion

Use the stated performance measure

Logic is the only rational method

Uncertainty and paths need other tools

Use probability, search or logic as appropriate

The program's internal score defines success

The task specifies success

Begin with the external performance measure

Autonomy means no built-in knowledge

Prior knowledge is permitted

Learn increasingly from experience

Later information should judge an earlier action

That is hindsight

Use information available when acting

Keep the courier quantities separate. P(W) is 0.34, not 0.18. P(blocked|W) is 0.18/0.34=9/17, not 0.90. The negative outcome contributes (9/17)(-6)=-54/17, making the total 90/17, about 5.29, below 8.

PEAS has its own traps. Sensors and actuators define the interface, the performance measure defines success, and the environment defines the task. A camera is not a percept sequence, a route is not an actuator, and expected score is not a guarantee. Under bounded rationality, an agent may seek the best feasible decision within time and computation limits.

Intelligent and rational agents: how questions test the foundations

Questions may ask you to identify a sensor, percept or actuator; construct PEAS; spot a changed rationality determinant; classify an environment; calculate a score or posterior; or judge claims about omniscience, autonomy and coordination.

Test yourself immediately:

  1. An unlikely bad outcome follows the best expected action. Was the decision necessarily irrational? No, if it used the available evidence and correct performance measure.

  2. Which route is selected before W? Short route, because 13.2>8.

  3. What is P(blocked|W)? 9/17, about 0.5294.

  4. If A has the lower robot ID, what bridge result follows? (Enter,Yield) with (4,1).

For a decision answer, write the performance measure, percept history, knowledge and actions. Calculate expected outcomes before naming the choice. For classification, state the assumption supporting each label.

Intelligent and rational agents: the short version and next step

Remember six anchors: an agent perceives and acts; f: P* -> A maps percept histories to actions; PEAS specifies the task; rationality depends on performance measure, percept sequence, prior knowledge and actions; uncertainty needs beliefs and expected outcomes; and rational agents may still need coordination.

For a mastery check, reconstruct the courier PEAS table, derive P(W)=0.34, recover 9/17, recompute EU(S|W)=90/17, and explain why a later-clear route does not make the detour irrational. Then reproduce the bridge matrix and explain how the priority rule changes the joint result.

The NTA-UGC-NET Paper 2 course provides a structured route whose live curriculum includes Artificial Intelligence and Multi Agent Systems. For a broader applied route, use AI & ML for Placements.