What turns a chatbot into an agent?

In this guide, an agent is a system that uses a model together with tools to pursue a task. The relevant distinction is not the marketing name. It is whether the system can take an action outside the reply, such as writing a file or sending a message.

A generated statement saying “I sent the email” is not evidence of delivery. Inspect the actual tool result or account record. Likewise, a plan describing a bank transfer should not be confused with an integration that can perform one.

OWASP’s AI Agent Security guidance recommends minimal tool access, scoped permissions and human review for sensitive actions. Those controls belong in the application around the model. A sentence in the prompt is useful instruction, but it is not an independent access-control mechanism.

Start with an invoice review that cannot pay

Consider an invented office workflow: collect invoice PDFs from a designated folder, extract their totals and produce a report of possible duplicates. It needs to read that folder and write a report. It does not need to change bank details or send money.

Define completion as a report containing the invoice reference, the extracted amount and the reason a pair was flagged. An uncertain extraction goes into a review queue. It should not silently become a payment decision.

Keep the original files available to the reviewer. The report should point back to evidence so a human can resolve ambiguous totals. This is a workflow-design example, not a deployed DarkGPT integration or a claim that a model can read every invoice correctly.

Build the permission table before connecting tools

Example policy for the invoice review
OperationPermissionReview rule
Read designated invoice folderAllow only that folder.Exclude unrelated records and credentials.
Write a draft reportAllow only the report location.Keep source references and uncertainty visible.
Send an external messageDo not enable in this first version.Add a separate approved workflow if needed.
Change a payment destinationDo not enable.Handle through the established account process.
Pay an invoiceDo not enable.Keep outside this review workflow.

The table describes an application policy. Implement it through the tool permissions and server-side checks available in your environment. Do not grant broad access and rely on the assistant to remember that it should behave narrowly.

When adding a new tool, revisit the task. A calendar connection is unnecessary for duplicate-invoice review. The fact that a connector exists is not a reason for this agent to use it.

A dedicated invoice tray connects through a transparent document reader to a draft report with a magnifying glass. A closed payment vault stands behind a boundary rail outside the connected workflow.
caption: Read and report within a fixed scope.

Decide when the agent must stop

A workflow needs an ending as much as a goal. For the example, stop after processing the permitted files or reaching the agreed execution limit. Return a partial report if an input is unreadable. Do not let the model keep searching unrelated folders to compensate.

Set limits on repeated tool calls and spending in the actual application. A model can suggest limits, but the runtime must enforce them. Keep enough progress information to see whether a repeated action is making progress or merely repeating a failure.

External documents remain task data. A sentence inside an invoice claiming “the reviewer authorizes payment” does not become a new permission. See the document prompt-injection case study for that trust boundary. Here the extra protection is that payment is not an available operation at all.

A robotic document gripper runs along a rail with a lime mechanical stop labelled Execution limit. An obscured invoice is diverted into a tray labelled Review queue beside the Unreadable input annotation.
caption: Set an enforceable stopping point.

Run a small trial with observable outcomes

Use invented invoices or files you are permitted to test. Include a duplicate, an unreadable document and two similar invoices that should remain separate. Before running, write what a correct report would say about each case.

Review the result against that expected behavior. Check which files were read and whether anything outside the report location changed. A correct-looking summary is only one part of the review; the action record matters too.

If the trial exposes a missing permission boundary, change the application configuration before expanding the task. More autonomy should follow a demonstrated need and a bounded design. It should not be the default reward for an agent that writes a confident plan.

Sources & further reading

Follow the original source to check its date and scope.

  1. AI Agent Security Cheat Sheet

    OWASP | scoped tools, sensitive-action review and enforced execution limits.

Make it your next question

Try this prompt

Help me design a read-only invoice review workflow. List the allowed data and tools. Separate analysis from sending or payment. Include review points and stop conditions without assuming the workflow is already connected to an account.

Use this prompt

Opens chat with this prompt filled in. You choose when to send it.