AI agent architecture is the set of layers that turn a language model into something that can complete work: the model itself, retrieval of knowledge, tools that take actions, memory of the conversation and customer, guardrails that constrain behavior, a handoff to people, and evaluation after the fact. A model on its own can only produce text. The surrounding layers decide whether that text is accurate, whether it leads to an action, and whether the organization can trust and review the result.
Understanding these layers helps buyers ask better questions and helps operators diagnose problems when an agent underperforms. This article describes each layer in plain language, explains how Bund AI implements it according to its documentation, and then turns to the AI agent use cases in support and sales where the architecture earns its value.
In brief: reliable agents are produced by the design around the model more than by the model alone. Retrieval and guardrails determine accuracy, tools determine usefulness, and handoff and evaluation determine whether the system can be run responsibly.
What are the main layers of an AI agent architecture?
Most production agents share seven layers. The model interprets the request and composes the reply. Retrieval supplies the facts. Tools let the agent act on other systems. Memory carries context across turns and channels. Guardrails check behavior and require approval where warranted. Handoff passes difficult cases to people. Evaluation measures what happened so the system can be improved.
These layers interact. A strong model with weak retrieval will answer fluently from the wrong source. A well-grounded agent without approval controls may take an action it should not. An agent without evaluation will degrade quietly as products and policies change. The sections below take each layer in turn. For readers choosing between products, the AI agent platform buyer's guide turns these layers into evaluation criteria.
How does the model layer work?
The model layer is the language model that reads the request, decides what is being asked and writes the response. Different requests need different capability: a short factual question can be handled by a fast, inexpensive model, whereas a long, multi-part or image-bearing question benefits from a stronger one.
Bund AI uses model routing for this purpose. According to its documentation, a fast model handles simple asks and a stronger model is used for long, multi-part or image-bearing questions, with no extra cost on simple turns. For the customer this is invisible. For the organization it means the cost and speed of routine questions are not set by the hardest ones.
How does retrieval ground the answers?
Retrieval is the layer that finds the relevant facts in your own content and gives them to the model. Without it, the model answers from general training data, which does not include your return policy or your delivery terms. Content is split into chunks, converted into numerical representations called embeddings, and stored so that the most relevant passages can be found by meaning rather than by exact wording. The approach and its pitfalls are covered in our article on retrieval that actually answers.
In Bund AI, knowledge can come from crawled URLs (re-crawled nightly), uploaded PDF, DOCX, Markdown or text files, pasted text and authored help articles. Each source is extracted, split into overlapping chunks, embedded and stored for semantic retrieval. When the agent cannot find an answer, it logs a knowledge gap so the team can supply the missing content. That gap log is a practical feature, because it converts the agent's uncertainty into a work list for the people who own the content.
What do tools and actions add?
Tools are what allow an agent to act rather than only to answer. The model decides that a request calls for a particular action, such as looking up an order, and the system executes it against a real system and returns the result for the model to use in its reply. A single request may involve several tool calls in sequence, which is often described as a tool loop.
Bund AI includes more than 20 built-in actions. These cover looking up an order, shipment or appointment, cancelling, rescheduling, changing an address, filing a claim, booking meetings through Google Calendar, Cal.com or Calendly, capturing leads and creating tickets. For anything else, custom HTTP actions let the agent call a business's own endpoints, with encrypted headers for authentication and an option to require approval for anything that writes. The self-service actions page describes how this works for refunds, changes and cancellations.
How should memory be understood?
Memory has two practical meanings. Within a conversation, the agent needs the earlier turns so that it does not ask the same question twice. Across conversations, it helps to know who the customer is and what has already happened. In Bund AI, the web widget and the email inbox share the same agent, memory, knowledge and tools, so a customer can be recognized across channels, and a rolled-up customer profile brings together chats, leads and tickets for each person.
Organizations should be deliberate about what memory retains. Personal data held across conversations carries privacy implications, and retention should follow the organization's own policy. The architecture decision to note is whether memory is shared across channels, since separate memories per channel produce inconsistent answers.
Which guardrails matter, and where does a human come in?
Guardrails operate at three points: before an action, around a reply and in the surrounding infrastructure. Before an action, a confirmation gate can pause anything that changes a real system. In Bund AI, sensitive actions can require one-tap approval on the web or a reply of YES by email. Around a reply, a compliance supervisor checks each answer for off-topic content, false claims that an action was completed, and ungrounded facts. In the infrastructure, card numbers are masked at storage, outbound fetches are blocked from reaching private addresses, and each workspace's records are isolated.
Handoff is the human layer. Bund AI escalates for clear reasons: low confidence in the answer, an explicit request for a person, a sensitive action that needs review, a supervisor block or a tool error. The escalation carries the full transcript and the customer's contact, so the person who takes over does not start cold. Our article on when AI should hand off discusses where the line should sit, and the escalation and handoff page shows the workflow.
How is an agent evaluated after launch?
Evaluation is the layer most often neglected. It answers whether the agent resolves what it is asked, where it fails and what it costs. Useful evidence includes conversation summaries, sentiment and customer satisfaction scores, the log of unanswered questions, escalation reasons and resolution rates over time. A safe test environment helps before and after changes: Bund AI offers a Playground in which side-effecting actions are simulated rather than run.
Bund AI also detects account signals such as churn risk, buying intent and complaints, attaches the customer's own words as evidence, and sends a weekly digest. For leaders, this turns the agent into a source of operational insight in addition to a channel; the customer insights page describes it. The company's headline figures, 86.7 percent of tickets resolved with no human touch and a 1.8 second median first response, come from early rollouts and should be read as indicative rather than as guarantees.
What are the main AI agent use cases in support and sales?
In support, the strongest use cases are repetitive requests where the answer and the action are both well defined. Examples include order status and tracking, returns and refunds within policy, address changes before dispatch, appointment changes and policy questions answered from documents. Each exercises several layers: retrieval for the policy, a tool for the action, a confirmation gate for the refund and a handoff when the case is unusual. The support automation and email support pages show these workflows, and AI customer support guide covers the surrounding practice.
In sales, the main use cases are answering inbound questions at any hour, asking the qualifying questions the business defines, detecting buying intent from the conversation, capturing the lead with context and booking a call against real availability. The lead capture and appointment booking pages describe the details. Industry variation matters too: an ecommerce store, a clinic and an insurer need different actions, which is why Bund AI ships tool packs for ten verticals.
What to consider before you rely on an agent
An architecture is only as good as its content and its policies. Retrieval cannot correct an outdated document, and guardrails cannot decide a policy the organization has not written. Teams should assign ownership of content, define which actions require approval, and review a sample of conversations regularly.
Bund AI is a closed-source, fully hosted product for sales and support, and its channels are the web widget and an email inbox. It is not the right choice for organizations that need to modify the architecture itself, self-host, or serve customers by telephone, SMS or WhatsApp. Those organizations may be better served by a framework or a broader platform; the no code AI agent builder article discusses that trade-off.
Frequently asked questions
What is AI agent architecture? AI agent architecture is the arrangement of components that lets a language model complete tasks: the model, retrieval of knowledge, tools for actions, memory, guardrails, human handoff and evaluation. The surrounding layers largely determine accuracy and safety, not the model alone.
What is the difference between retrieval and memory in an AI agent? Retrieval finds facts in the organization's own content, such as policies and product information, so answers are grounded. Memory carries the context of the conversation and the customer, so the agent does not repeat questions and can recognize returning customers.
How does an AI agent take actions? The model decides that a request needs an action, and the system runs a tool, such as an order lookup or a reschedule, against a real system and returns the result. Sensitive actions can be paused for human approval before they run.
What are common AI agent use cases? In support, common use cases are order tracking, refunds within policy, address and appointment changes and policy questions. In sales they are answering inbound questions, qualifying leads, capturing contact details and booking meetings.
How does Bund AI handle handoff to a human? Bund AI escalates on low confidence, an explicit request for a person, a sensitive action, a supervisor block or a tool error. The escalation carries the full transcript and customer contact and can be auto-assigned to a teammate.
Is Bund AI open source? No. Bund AI is a closed-source, fully hosted product, so the architecture described here is configured through the product rather than modified at source level.