What you will build understanding of
The companion animated episode runs 24 minutes and 28 seconds. The chapters below follow its teaching sequence and include full-size diagrams. The video uses original animation and stock synthetic narration.
The fictional Cedar Workshop examples use order O17, customer C104 and policy KB-01. They demonstrate how an application controls execution. The Python pack uses recorded model decisions so you can run it without an API account.
Start with the job

Build effective AI agents.
An impressive agent demo can hide a very ordinary engineering problem. A customer asks where an order is. Somewhere, a system needs to find the right order, read its status, check a delivery policy, and give a useful answer. The hard part is making that sequence work when information is missing, a service times out, or the request is unusual. In this episode, we will build a clear mental model for effective AI agents, and the simpler workflows that often solve the same job.
The goal: an answer you can support
We will follow one fictional workshop through the explanation. Its assistant handles order questions and prepares draft quotes. That gives us a concrete way to examine each pattern, instead of collecting names for architectures. We will see where the model adds value, where ordinary code should take over, and where a person needs to make a decision. The animations show information moving through the system. They are explanations, not recordings of a live model making those decisions. The companion examples use clearly labelled, deterministic test fixtures.
Workflows and agents

Who chooses the next step?
First, separate two meanings that often get mixed together. In a workflow, the application defines the path. A model might classify a message or write a response, but your code decides what happens next. In an agent, the model can choose the next action based on what it has learned so far. It might search, inspect an order, ask a follow-up question, or finish. This distinction concerns control flow. A system does not become an agent simply because one of its functions calls a language model.
A workflow can still use an LLM
Imagine a delivery question arriving at the workshop. Our workflow always classifies it, looks up the order, retrieves the relevant delivery information, and prepares an answer for review. The model can interpret natural language at the first step and explain the result at the last. The steps between those calls remain explicit. We know which service will be called and what fields it must return. That makes failures easier to locate. If the order is missing, we can stop at that point and ask for the order number.
An agent chooses the path as it goes
Now change the request. A customer says the order is late, the address changed, and the tracking page disagrees with an email. We may not know the right sequence of investigations in advance. An agent could inspect the order, compare the tracking record, and then decide that it needs the earlier conversation. Its next step depends on the previous observation. That flexibility can be useful. It also creates more possibilities to test. Your application still decides which tools exist, which actions need approval, and when the agent must stop.
Choose control before complexity
Neither architecture is automatically better. A fixed process can be the right answer for a frequent, well-defined request. A flexible process can help when the task genuinely changes as new information arrives. Start by writing down a successful outcome and a few failure cases. Then ask whether the steps are already known. If they are, an explicit workflow gives you a useful baseline. If they are not, experiment with a bounded agent and compare it against that baseline. The name of the architecture is less useful than the behavior you can demonstrate.
Pick your building tools

Begin with the smallest useful system
Before choosing a framework, ask what actually requires a model. Multiplying a quantity by a trusted price does not. Reading an exact order identifier from a database does not. Understanding an ambiguous message may. Drafting a clear explanation from several records may. A useful first version often combines ordinary application code with one carefully specified model call. Add retrieval if it needs external knowledge. Add a workflow if the job has separate stages. Add autonomy when a fixed sequence is demonstrably insufficient. Each step should solve a problem you have observed.
Code and visual builders share the same responsibilities
You can implement these ideas in Python, TypeScript, or a visual workflow builder. The interface changes, but the responsibilities remain. Inputs need a defined shape. Branches need clear conditions. Failures need a destination. Credentials need appropriate handling. A visual canvas does not remove those responsibilities, and writing code does not automatically satisfy them. Choose a tool that lets you inspect the request, the prompt, the tool call, and the result. If a step fails, you should be able to explain which part failed without guessing what a hidden layer did.
Understand one complete request
Frameworks can save time on message handling, tool registration, persistence, and tracing. Use that help when it fits the application. But first understand one complete request without the framework's abstractions. What exactly did you send to the model? What came back? Who executed the tool? How was the result passed into the next call? That understanding is what lets you debug a surprising answer. It also makes it easier to change providers or frameworks later, because the behavior you depend on is written down and tested.
The augmented LLM

Give the model the context the task needs
The first building block is a model connected to useful context. Three common additions are retrieval, tools, and memory. Retrieval supplies relevant information from outside the conversation. Tools let the application perform an operation or fetch current data. Memory preserves useful information from earlier interactions. They can work together, but they solve different problems. For our workshop assistant, the delivery policy comes from retrieval, the order status comes from a tool, and the customer's earlier clarification comes from the conversation history.
Retrieval finds evidence
Suppose the customer asks how long delivery takes after workshop completion. The model should receive the relevant policy, rather than invent a delivery window from general knowledge. A retrieval system searches the available material and returns candidate passages. That search might use keywords, embeddings, or a combination. The model then answers using the selected evidence. Retrieval is not a guarantee that the right passage was found. If the database contains an outdated policy and a current one, the system needs a way to prefer the current source and expose uncertainty when the evidence is inadequate.
Tools fetch facts or perform actions
A tool is an interface your application can call. For example, look up order O seventeen and return its current status. The model may request that call, but application code executes it. That separation matters. A proposed tool call is not proof that the operation happened. The assistant should wait for the actual result before claiming success. Read-only tools are useful starting points because they let you learn about tool selection and argument validation without changing external records. Tools that send messages or modify data need clearly defined permissions and review points.
Memory is selected context
Memory can be as simple as the messages in the current conversation. If a customer already provided an order number, the next response should not ask for it again. But a growing history is not automatically a better history. Old details can conflict with newer corrections, and unrelated content consumes space. Keep the information that helps the current task, preserve the source of important facts, and make changes understandable. A saved preference, a retrieved policy, and a current database record should not be treated as interchangeable evidence.
Combine context without confusing its sources
Put the three additions together. The conversation tells us which order the customer means. A tool reports that workshop work is complete. Retrieval returns a policy saying standard delivery takes three business days after completion. The assistant can now prepare a grounded response. It should not turn that policy into a guaranteed arrival date if no shipping estimate is available. This is a small but important distinction: giving the model more information helps only when the system also preserves what each piece of information actually means.
Prompt chaining

Pattern 1 · Prompt chaining
Prompt chaining breaks a task into a sequence of smaller stages. Each stage receives a defined input and produces something the next stage can use. Consider a draft quote. One model call extracts the requested service and quantity from a message. Code checks those fields. Code then calculates the amount from a trusted catalog. A later model call explains the draft in clear language. We do not ask a single prompt to interpret the request, invent missing details, calculate prices, and send the result. We make those responsibilities visible.
A gate stops a bad intermediate result
The important part of a chain is the boundary between stages. If the extracted request has no quantity, the price calculation should not continue with a guessed value. Our example returns an ask-user result and identifies the missing field. If the customer identifier is unknown, it goes to review. These are programmatic gates: ordinary checks that decide whether an intermediate result is usable. A model can help interpret a response, but it should not overrule a required field just because proceeding would make the conversation sound smoother.
Trace the successful path
Now the customer confirms two inspections. The trusted catalog contains a price of forty-five dirhams per inspection. The calculation returns ninety dirhams, and the result is explicitly marked as a draft. The explanation step can describe the service and amount, while keeping the review requirement. In the companion example, these values come from executed application code. The language-model extraction is represented by a fixture so you can inspect the control flow without an API key. Passing that test proves the gate and calculation behavior, not the accuracy of a live model.
Use a chain when the stages are known
Chaining helps when the problem separates into stable stages that you can inspect. It does not guarantee that every extra model call improves the answer. More calls can add latency, cost, and new failure points. Measure whether the decomposition helps on examples that represent your actual workload. Keep calculations, formatting rules, and exact lookups in code where that is appropriate. Use model calls for the parts that benefit from language understanding or generation. A good chain gives each stage a clear job and makes a bad handoff visible early.
Routing

Pattern 2 · Routing
Routing becomes useful when incoming requests need different handling. A delivery question should use delivery information. A technical problem should enter a troubleshooting process. A billing dispute may need a person to inspect the account. A classifier proposes a category, and application code selects the corresponding path. The classifier might be a language model, a conventional model, or a set of rules. The key idea is that different categories get appropriate handling instead of forcing every request through one enormous prompt.
Unknown categories need a destination
Treat the route as structured data, not a sentence that your application tries to interpret loosely. Define the allowed categories. Check that the output belongs to that set. Decide what happens if the category is unknown or ambiguous. In our demonstration, an unfamiliar category goes to human triage. A request must not gain access to a powerful operation simply by producing a convincing label. The router selects among routes that the application already permits. It does not create new permissions.
A confidence number needs evidence
The example also includes a confidence threshold to illustrate an abstention path. Those numbers are synthetic fixtures, not measured model confidence. In a real system, a model saying it is ninety-six percent confident does not make that probability calibrated. You need evaluation data to understand what a score means and where to set the threshold. Review confused categories, especially those with different consequences. It may be acceptable to ask an extra clarification question. Sending a billing complaint into an automatic delivery reply can create a much worse experience.
Automate one category first
A practical rollout starts with a narrow category. For example, automate straightforward order-status questions and send everything else to the existing team. That gives you a defined slice of traffic to evaluate. Track whether the classifier selects the right requests, whether order lookup succeeds, and whether the final draft uses the evidence correctly. When that slice performs well enough for your requirements, add another category. This approach gives you a way to expand based on observed behavior instead of assuming that success on a demo covers every customer request.
Parallelization

Pattern 3 · Parallelization
Parallelization runs independent work at the same time. Suppose a draft needs several checks: does it refer to the right customer, does it include supporting evidence, and does it falsely claim that a message was already sent? Those checks can inspect the same draft independently. We can run them concurrently and combine their results afterward. If a later task requires the output of an earlier one, those tasks are not independent and should not be launched as though they were. The dependency structure determines what can run together.
Concurrent work can reduce waiting
The timing diagram shows the difference. Sequential checks wait for one another. Independent concurrent checks can finish closer to the duration of the slowest check, plus coordination overhead. That is an architectural illustration, not a measured speedup for every provider. Rate limits, resource contention, and slow dependencies can change the result. You also need a policy for partial failure. If an essential check times out, the system should not silently treat the missing result as a pass. Required evidence must remain required when a service is slow.
Agreement is not proof
Another use of parallel calls is to request several judgments and combine them. That can reveal disagreements, but agreement does not prove correctness. Several model calls can share the same mistaken assumption or missing context. Use independent evidence and objective checks where possible. In our companion example, the three checks are simple Python predicates, so their behavior is easy to inspect. In a real application, some judgments may require a model or a person. Choose the method according to the question being checked, and evaluate the combined decision.
Orchestrator and workers

Pattern 4 · Orchestrator and workers
An orchestrator decides which subtasks are needed for a particular request. Workers carry out those tasks, and their results are combined. This is useful when the list of subtasks changes with the input. A simple delivery question may need one lookup. A contradictory shipping question may require an order record, a carrier update, and a policy search. The plan is generated from the situation rather than being identical for every request. That flexibility is the main difference from a fixed set of parallel checks.
A generated plan still needs validation
Do not execute a plan merely because it looks sensible. Check that every requested operation exists and is permitted. Limit how many tasks can be created. Validate dependencies so that a worker does not wait on a result that will never exist. Our example permits an order lookup, a policy search, and a draft response. The draft depends on the first two results. An invented refund operation is rejected. These checks constrain the plan without pretending that the application already knows every useful sequence of investigations.
Dynamic plans can run sequentially or concurrently
The workers do not have to be strictly sequential. Independent lookups can run concurrently, while a draft waits for the facts it needs. What matters is that the plan captures those dependencies. Keep the results associated with their source and task identifier so the final synthesis can explain its evidence. When a worker fails, report that failure to the orchestrator or to a person. Avoid turning a partial collection of results into an apparently complete answer. The overall workflow is only as trustworthy as the information it actually obtained.
Evaluator and optimizer

Pattern 5 · Evaluator and optimizer
An evaluator-optimizer loop creates an output, checks it against criteria, and revises it using feedback. For our delivery response, the criteria are concrete: state the policy accurately, include its evidence identifier, and avoid an unsupported guarantee. The first draft says that the parcel will arrive soon. That sounds friendly, but it omits the useful information. The evaluator identifies what is missing. The next attempt can fix those specific issues. This is more actionable than repeatedly asking a model to make the answer better without saying what better means.
Feedback should change a specific defect
Watch the revision. The vague draft becomes a statement that standard delivery takes three business days after completion, with the source identifier attached. The change addresses the two failed criteria. That does not mean the sentence is now universally correct. It is correct only if this is the relevant current policy and the underlying order state supports it. Evaluation needs access to the evidence being checked. Otherwise, an evaluator can confidently approve a fluent statement while missing the same factual problem as the generator.
Stop when improvement is no longer useful
Put a limit on revisions. A loop that keeps rewriting without measurable improvement consumes time and makes the final result harder to predict. Our example allows three attempts and returns a review state if the criteria remain unmet. You can also compare revisions to detect repeated answers or regressions. A separate model call is one possible evaluator, but exact checks should remain exact checks. Use code for required fields and arithmetic. Use a model or human judgment for criteria that genuinely need interpretation.
The agent feedback loop

The agent loop: act, observe, decide
Now we can see what changes in an agent. A person gives it a goal and an environment with available tools. The model selects an action. The application checks and executes that action. The resulting observation goes back to the model, which chooses what to do next. The number and order of steps can vary. It might finish quickly, discover missing information, or ask for help. The loop should be driven by actual observations from the environment, not by the model assuming that its previous action succeeded.
Follow one agent run
In our workshop example, the first decision reads the order. The result says that work is complete. The next decision retrieves the delivery policy. The agent now has enough information to draft a response, but sending it is a separate action with a review requirement. The trace records each operation and its returned data. This makes the run inspectable. You can see which evidence supported the answer and where it stopped. Our recorded decisions are fixtures that demonstrate the loop mechanics; they are not presented as a benchmark of model reasoning.
A small loop can have a large behavior space
The basic control loop can be short. That does not make the resulting system easy to validate. Each decision creates branches, and tool failures create more. Set a maximum step count and a time budget. Define a finish condition. Decide which situations need human input. Restrict the available tools to the current task and user. A request to use an unknown tool should fail explicitly. A request to send a message should reach the application's permission boundary. These controls belong in the application, even when the model is choosing the path.
Coding agents make the loop easy to see
A coding assistant is a useful illustration. It inspects a project, changes code, runs tests, reads a failure, and revises the change. The test output provides feedback from the environment. Passing tests are useful evidence, but they cover only the behavior that was tested. A change can still violate an unstated requirement or break an untested path. That is why a successful agent run includes reviewable changes and relevant validation. The result should be judged by what it did and what was verified, not by how confident its final explanation sounds.
Use autonomy where the path must adapt
Agents are most useful when adapting the next step has real value and the environment provides feedback you can evaluate. If the same four steps solve the task every time, a workflow may be easier to operate. If an investigation genuinely branches as facts arrive, an agent may help. Compare both using representative cases, including failures. Do not turn a claim about one product in an old demonstration into a claim about every current agent. Capabilities change; the engineering question remains whether the system meets your requirements on the tasks you care about.
Make the system reliable

Build one reliable slice, then expand
The first production lesson is to narrow the problem. Choose a request category you understand well enough to describe how a person solves it. Collect representative cases, including confusing and incomplete ones. Define what success looks like and what should be escalated. For the workshop, start with order status before adding billing disputes or changes to orders. A narrow scope gives you a clearer view of failure causes. It also helps you decide whether the missing piece is better retrieval, a clearer tool, a different prompt, or a more flexible architecture.
A small error rate becomes visible at scale
Scaling changes what you notice. A rare failure may never appear during a short demonstration, then occur regularly when many people use the application. More data can also make retrieval harder by adding duplicates, contradictions, and irrelevant passages. Monitor the kinds of cases that arrive and the sources that are being selected. Keep a way to route uncertain cases to people. Launching to more users should follow evidence about quality and operational capacity, because successful handling of a few hand-picked examples does not establish how the system behaves across real traffic.
Would a prompt change improve the system?
Evaluation gives you a way to answer that question. Keep a set of cases with expected outcomes or explicit scoring criteria. Run the old and new versions on comparable inputs. Look at factual correctness, task completion, unnecessary escalation, latency, and cost where relevant. Inspect regressions rather than hiding them inside one average score. Model behavior can vary between runs, so a tiny improvement on a tiny sample may be noise. The goal is a repeatable decision process for changing the system, not a single number that makes the dashboard look good.
Control-flow tests and model evaluations are different
The examples behind this video pass twenty application checks. They verify things such as rejecting an unknown route, stopping after a step budget, detecting missing evidence, and requiring review before a send action. That is valuable, but it is not an evaluation of a live language model. To evaluate the model, we would need representative prompts, real model responses, scoring criteria, and repeated comparisons. Keep those claims separate. A tested application can still receive poor model decisions, and a capable model still needs correctly implemented application controls.
Checks should match the failure you want to prevent
Use checks at meaningful boundaries. Validate the inputs before calling a tool. Compare generated identifiers with trusted records. Check that an answer's sources support its claims. Enforce permissions before changing data or sending a message. A second model can help judge language or relevance, but it is not a universal safety guarantee. Retrieved documents and tool outputs can contain misleading instructions, so treat them as data rather than authority to change the task. The purpose of each check should be specific enough that you can test what happens when it fails.
Make failures explainable
Finally, make the system observable. Record enough information to explain which route was chosen, which tools were called, what they returned, and why the run stopped. Keep sensitive content out of logs unless it is actually needed and appropriately handled. A useful trace lets an engineer distinguish a retrieval failure from a routing mistake or a timeout. It also gives the team real examples to improve. Without that visibility, every bad answer looks like a vague model problem, and the response tends to be another prompt change without a clear diagnosis.
Choose your next step

The practical sequence
Here is the practical sequence. Define the job and the acceptable outcomes. Build the smallest useful baseline. Add retrieval, tools, or memory when the task needs them. Use chaining for known stages, routing for different request categories, and parallel work for independent tasks. Use an orchestrator when the subtasks vary, and an evaluator loop when specific feedback can improve an output. Consider an agent when the path must adapt to new observations. At every stage, make the behavior inspectable and compare it with representative cases.
Build something you can explain and verify
You now have a way to choose an architecture based on the work it needs to do. Start with one useful problem. Draw the inputs, the decisions, the evidence, and the stopping conditions. Then build and test that path before expanding it. The accompanying guide includes the diagrams, worked examples, and Python fixtures at with Hammad dot com. The source references are listed there as well. Save this episode as a reference, and use the patterns to build a system whose results you can explain and verify.
Scope your first workflow
Copy this prompt and replace the bracketed details. Use real examples you are permitted to share. Treat the proposed plan as a starting point for review.
Help me scope one useful AI workflow. Task: [one recurring job] Users and permitted data: [who and what] Inputs: [examples, including incomplete requests] Expected output: [a result someone can verify] Available tools: [read and write operations] 1. Describe a simple baseline and how to judge success. 2. Identify which steps can be ordinary code. 3. Choose a workflow pattern only where it helps. Explain the tradeoff. 4. List sources, input checks, permissions, failure paths and stopping conditions. 5. Propose representative evaluation cases and separate model-quality evaluations from application tests. 6. Return a diagram description and the smallest implementation plan. Flag unknowns; do not invent results.
Run and inspect the examples
The pack includes chaining, routing, parallel checks, plan validation, evaluator feedback and a bounded tool loop. It requires Python 3.10 or newer and no extra packages.
python3 -m unittest -v test_patterns.py python3 demo.py
All 20 application tests passed during production. They check defined program behavior; they do not establish a live model's accuracy, production readiness or business results.
Python examples and 20 tests ↓Planning prompt ↓Complete written guide ↓Recorded test results ↓Sources and further reading
- Anthropic: Building effective agents — primary reference for the workflow patterns and the distinction between workflows and agents.
- Dave Ebbelaar: How to Build Effective AI Agents (without the hype) — the requested video reference, published in January 2025.
This independent lesson uses original explanations, fictional examples and diagrams. The reference creator's script, footage and audio are not redistributed. Historical product anecdotes are not presented as current benchmarks.