What is an AI agent, and where do businesses use one? A guide

An open brass clock movement beside a pocket watch with its case open, a pair of tweezers and a jeweller's loupe

An AI agent is software that breaks a goal into steps on its own, uses tools to carry out those steps, and checks the result before moving on. The difference from a chat model that simply answers questions is that an agent does things — it looks up a record, reads a document, fills in a form, starts a process. In businesses, agents are used most for triaging customer requests, document and data work, running internal tools, and monitoring.

Here we explain what an agent is, how it works, where it helps and where it does not, and what we pay attention to when we build one.

What is an AI agent? A short definition

An AI agent has four parts:

  • A model. Usually a large language model. This is the “mind” that decides what to do.
  • Tools. The places where the agent touches the world: a database query, an API call, sending an email, reading a file.
  • A goal and instructions. The agent is told what it must achieve and which rules it must stay within.
  • A loop. The agent takes a step, looks at the result, then decides on the next step. It carries on until it reaches the goal or hits a limit.

An analogy: a chat model is a well-read adviser who answers whatever you ask. An agent is what you get when you give that same adviser a desk, a telephone and access to a few systems. Now they do not only tell you — they act.

Chatbot, automation and agent: where is the difference?

These three are often confused. A short comparison:

Classic automationChatbotAI agent
How does it decide?By pre-written rulesGenerates an answer to the questionChooses steps towards a goal
Does it use tools?Yes, in a fixed orderUsually notYes, depending on the situation
When something unexpected happensStops or errorsGives a general answerTries another route
How easy to audit?EasyModerateDepends on the design

The last row matters. An agent’s flexibility is also its greatest risk. A rule-based automation does the same thing every time. An agent does whatever it judges best each time. That is why the boundaries must be drawn at the very beginning.

How does an agent work? A concrete example

Take an online shop. A customer writes: “The order I placed last week still hasn’t arrived.”

A rule-based automation looks for a keyword and sends a canned reply. An agent might proceed like this:

  1. It finds the order record from the customer’s phone number.
  2. It queries the courier’s system for the shipping status.
  3. It sees that the parcel is waiting at a sorting hub.
  4. It writes to the customer with an estimated delivery date and a tracking link.
  5. If the delay exceeds a set threshold, it hands the request to the customer service team, with its reasoning attached.

The agent made five separate decisions in that flow. None of them, though, was a step that is hard to undo — such as a refund or a cancellation. That is a deliberate choice. What an agent may do on its own, and what it may only suggest, is written down in the design.

Where do businesses use AI agents?

In our experience, these are the areas where agents genuinely earn their keep.

Triaging requests

Reading incoming emails, forms or messages; classifying them; asking for missing information; routing them to the right person. The decision stays with a human — the agent puts a tidy, summarised file in front of them.

Document and data work

Reading invoices, contracts or applications, extracting specific fields and flagging inconsistencies. Assistants that work with company documents belong here too; we explain how they are built in what is RAG.

Running internal tools

An agent that prepares a report when a team member says, “Pull last month’s returns and tell me the most common reason.” Here the agent connects securely to the company’s own systems. MCP (Model Context Protocol) is one of the protocols that standardise that connection.

Monitoring and alerts

Watching a continuous stream of data — orders, system logs, market data — and raising an alert when something unusual happens. Our own autonomous trading system, CAI, is an extreme case of this kind of monitoring and decision process: ten agents that continuously read the market regime and manage the risk budget in real time.

One agent or many?

Most business projects start with a single agent, and work perfectly well with one. When the task is narrow and the tools are few, splitting the agent adds needless complexity.

But when a decision depends on combining different kinds of information, when conditions change often, or when a mistake is costly, dividing the work between several specialist agents builds a sturdier structure. An orchestrator coordinates them. We explain this approach, and when it is needed, in what is a multi-agent AI system, using CAI as the example.

The limits of agents: an honest list

Agents are powerful, but they are not magic. Before starting a project, it helps to know:

  • They can be wrong. Language models sometimes produce false information with complete confidence. If an agent acts on it, the error grows.
  • Cost grows with the number of steps. Every loop and every tool call is another model call. An agent left without limits can produce an unexpected bill.
  • Permission is risk. Every permission you grant an agent is a permission that can be misused. The difference between “can read” and “can delete” must be explicit in the design.
  • Not every job needs an agent. For a process whose steps never change, a rule-based automation is cheaper, faster and more dependable.

What we pay attention to when building an agent safely

There are a few rules we write down at the start of every agent project:

  1. A narrow task. Not “run customer service”, but “answer questions about delayed deliveries and hand over everything else”.
  2. A permissions list. Which tools the agent may use and which data it may access are written out one by one. The default is “no”.
  3. Human approval. Steps that are hard to reverse — payments, deletions, anything sent outside the company — need a person to approve them.
  4. Step and spending limits. If the agent cannot reach the goal within a set number of steps, it stops and asks a human.
  5. Logging. Every step and its reasoning are recorded. “Why did it do that?” always has an answer.
  6. Measurement. We start with a small prototype and measure accuracy, cost and speed on real examples.

Most of these rules are about design, not technology. That is why our agent projects usually begin with a short written architecture note.

Frequently asked questions

Is an AI agent the same as a chatbot?

No. A chatbot talks; an agent talks and also gets work done. A chatbot can be built as an agent, but not every chatbot is one. The difference lies in using tools and choosing steps.

Can a small business use an AI agent?

Yes. In fact the best results often come from small, narrow tasks: classifying incoming requests, extracting information from documents, preparing a weekly report. What matters is choosing the task well.

Will the agent send our data outside the company?

That depends entirely on the design. Which data goes to which model, where it is processed and where it is stored are set out in writing at the start of the project. For sensitive data, that decision is the first item in the architecture.

What do we need to build an agent?

A narrow, clearly defined task, a list of the systems the agent will reach, and an idea of how success will be measured are enough to begin. Real examples — past requests, documents, records — help a great deal in both design and testing.

What happens if the agent makes a wrong decision?

We build the system with that in mind. Critical steps require human approval, the agent stops when limits are crossed, and every step is logged with its reasoning. That keeps mistakes small, and their causes visible.

If you would like to talk about where an agent could make a real difference in your business, see our AI systems architecture page or write to us.

Open a conversation