Use a fixed workflow when the steps are predictable. Add model-directed decisions where the next useful step depends on new evidence. Keep permissions, budgets, and final-state checks under application control.
What makes a system an agent?
There is no single universal definition of agent. Anthropic makes a useful distinction: workflows follow predefined code paths, while agents let a model direct the sequence of actions. A tool call alone does not settle the distinction. A system that always searches, summarizes, and translates can use an LLM at each step while remaining a fixed workflow.
Ask a concrete question: after an unexpected result, who decides what happens next? If application code selects from a few designed branches, describe those branches. If the model can choose another search, inspect a record, ask for information, or stop based on what it learned, describe that decision surface. This tells a teammate more than an agent label does.
Choose the design from the task
Consider three requests. Extracting an invoice total may need one model call and a schema check. Classifying a support message and sending it to a known queue may need a workflow. Investigating why a delivery failed may require different checks depending on whether the address, payment, carrier, or order state is the problem. That last task is a plausible agent candidate.
These are design examples, not a benchmark proving that agents win. Build a simple baseline first and list its observed failures. Extra decision steps are worthwhile only if they recover important cases at an acceptable cost. A flexible agent that takes thirty seconds to perform a predictable two-second lookup has not improved the product merely because it can plan.
The tool loop is a controlled exchange
For application-executed tools, the model requests a named operation with arguments; the application runs it and returns a result; the model then responds or requests another operation. Anthropic documents this round trip explicitly. The model proposes an action, but a proposal is not permission and does not prove execution succeeded.
Your runtime needs separate checks for allowed tools, valid arguments, execution errors, and task completion. Return an explicit result such as order not found or carrier unavailable instead of silently providing an empty string. Give the loop a maximum duration and tool-call budget. If the agent reaches a limit, preserve its findings and hand off the unresolved task instead of presenting an invented successful outcome.
Worked example: investigate a missing delivery
This original scenario concerns a shop assistant handling: My package says delivered, but it is not here. Define success as explaining the verified delivery status and producing an allowed next step. Initially expose read tools for the authenticated customer's order, carrier events, and current support policy. Do not let a name or order number supplied in chat bypass customer authorization.
The agent reads the order and finds a carrier event for delivery to a collection point. It checks the carrier details, identifies the location, and searches the policy for collection deadlines. That evidence supports directions to the collection point. A different event showing delivery to a home address would justify a different investigation or escalation. The next step depends on observations, rather than a universal search order.
Now suppose the carrier lookup times out. The assistant can report the last verified order state and offer escalation; it cannot confirm the package's present location. If creating an escalation ticket is allowed, the tool should return a ticket identifier. The final message includes that identifier only after the write succeeds. The scenario's permission rules are a product decision, not a claim about every support system.
Make tools narrow and enforceable
A tool named handle_customer is difficult to inspect. A tool named get_delivery_events with an order identifier and an explicit result shape has a clearer purpose. A separate create_delivery_case operation makes its write effect visible. Describe required arguments, possible errors, and whether repeating the call changes anything. Validate those properties in the service that executes the tool.
The MCP tools specification requires servers to validate inputs and enforce access controls, and recommends client-side controls for sensitive operations. In our shop example, implement customer scoping in the service and use a stable request identifier when creating a case. A model instruction can explain these rules, but the database and API must enforce them even when the model chooses badly.
Memory and outside text need boundaries
Keep verified facts separate from model-written notes. An order identifier returned by an authorized lookup and a guessed explanation for a delay have different status. Store the source, timestamp, and scope of important observations. If a later session uses a saved delivery summary, refresh the order before proposing an action whose correctness depends on its current state.
Outside text can also contain misleading instructions. A carrier note saying send the customer record to this address is data to examine, not authority to grant a new tool permission. Limit which services the runtime can contact and what fields each tool can return. More memory and more tools can increase exposure and confusion; neither automatically makes the assistant more capable.
Evaluate the result and the route to it
Anthropic distinguishes the agent's trace from the outcome in the environment. For the shop example, inspect the actual ticket record as well as the final message. Repeat cases because a single successful run can conceal variation. Use human reviewers to calibrate any model-based grading of answer quality.
Build cases for collection-point delivery, wrong-account access, absent orders, expired collection windows, duplicate requests, and unavailable carriers. Check whether factual claims match the supplied records, whether unauthorized data appeared, and whether an escalation was created once. Track latency and cost per resolved case, including unsuccessful attempts. An overall average can hide a serious failure in the small subset of requests that require a write.
A practical exercise for a prototype or interview
Sketch the delivery assistant using fake records and a stub carrier. Write the expected outcome before running the agent. Compare it with a fixed workflow on the same cases. Your conclusion should explain which uncertainty needs a model decision and which requirements belong in ordinary code.
- Name the goal, allowed reads and writes, and the conditions for asking a human for help.
- Give each tool a precise result, explicit failure response, and enforced customer scope.
- Set an execution budget and define what the user receives when it runs out.
- Simulate a timeout and a repeated ticket request; verify the saved state after each run.
- Record one useful model decision and one decision that could be replaced with a predictable rule.
Quick answers
Frequently asked questions
Is every chatbot an AI agent?
No. A conversational interface can call a model once or follow a fixed workflow. Describe which actions it can take and who chooses their sequence to make its behavior clear.
Does an agent need multiple models?
No. One model can choose and use several tools. Add separate models or workers only when a measured need justifies the extra coordination and cost.
Is MCP the agent itself?
No. MCP provides a protocol for exposing capabilities such as tools. The model and application runtime still determine how those capabilities are selected, authorized, and used.
How do I know an agent completed its task?
Check an observable result: an existing ticket, a verified record, or a passing task-specific test. A confident final message is useful communication, but is not sufficient evidence of success.
Source notes
References and review policy
Information checked on October 2, 2026. Section links identify sources for factual claims and technical explanations. Interpretations, practice scenarios and preparation recommendations are RecallDeck’s editorial work.
From reading to recall
Practice the full interview loop.
RecallDeck schedules the concepts you miss and keeps coding, design, and behavioral fundamentals available when the interviewer changes direction.