Before building an AI agent, write down seven things: the task and its trigger, the rules and exceptions, what a person must approve, which systems it may touch, examples of good and bad output, the person who owns it, and how you will measure it. If you cannot fill those in, the build is premature. The documentation is also most of what makes the agent dependable.
First, what is an agent?
The word is used loosely. Anthropic's engineering guidance separates two designs: workflows, where models and tools follow predefined code paths, and agents, where the model directs its own steps and tool use. It advises starting with the simplest solution and adding complexity only when needed, because agents tend to trade cost and speed for flexibility and can compound errors (see sources).
For most small-business tasks, the right design is closer to a workflow with guardrails. We use "agent" on this site for AI-assisted work on a repeated task that a person supervises. Whatever you call it, the documentation below applies.
The seven things to write down
| Write down | The question it answers | Hypothetical example: reply drafts for quote requests |
|---|---|---|
| Task and trigger | What starts it, and what counts as done? | A quote request arrives by form. Done = a draft reply is waiting for review. |
| Rules and exceptions | What does a good result follow, and when must it stop? | Use the approved service list. If the request is outside it, flag for the owner instead of drafting. |
| Approvals | Which steps need a person, and who? | A person reads every draft. Anything mentioning price or dates is edited by the owner. |
| Systems and permissions | What may it read and change? | Read the inquiry. Write a draft note on the record. It cannot send messages. |
| Examples | What do good, bad and borderline outputs look like? | Ten past inquiries with the reply the owner would send, including two awkward ones. |
| Owner | Who is accountable, and who changes the rules? | The office manager owns it. The owner approves rule changes. |
| Measure | How will you know it is helping? | Time from inquiry to first reply; drafts accepted with little editing. |
Who approves and what the agent does
The most useful sentence in an agent specification is a plain statement of what is the agent's job and what is not. A table keeps it honest.
| Step | Agent | Person |
|---|---|---|
| Read the incoming request | Reads and summarizes it | Confirms the summary on review |
| Check against the approved service list | Applies the written rules | Decides every exception |
| Draft a reply | Writes it in the approved voice | Edits and approves before anything is sent |
| Mention price, dates or commitments | Does not. Flags the request | Decides and writes this part |
| Record what happened | Logs the draft and the decision | Reviews the log on a set rhythm |
| Change the rules | Cannot | The owner, after reviewing logged problems |
Test with real cases, including the awkward ones
Collect ten to twenty real past cases before the build, with the outcome you would have wanted. Include the ones that were messy: incomplete information, an angry customer, a request you do not offer. Run the agent over them and read every output. The aim is to learn where it fails and what it does when it fails, not to hit a score. Keep the set; you will rerun it whenever the rules or the model change.
Decide the failure paths in advance
- What happens when information is missing? (Ask a person, never guess.)
- What happens when a connected system is down? (Queue and alert the owner.)
- How does someone switch it off, and who is allowed to?
- Where are the logs, and who reads them?
Own the pieces that matter
Agree before the build who owns the accounts, the prompts, the workflow and the documentation. At QS Digital, accounts are opened in the client's name and the prompts, workflows and SOPs we write are the client's. The AI models and hosted platforms underneath are rented from their providers, so a good build documents what depends on a vendor. See what you receive and control.
Where a consultant helps and where you do not need one
You can write most of this yourself, and we would rather you did than paid for a build on a process nobody has described. Where help pays off is in finding the rules people apply without noticing, deciding approval points, building the test set, and wiring the agent into your systems safely. If that is the stage you are at, custom AI agent buildouts follow exactly this order, and a single paid hour can be used to review your draft documentation.
Sources and evidence
- Anthropic, Building effective agents. Primary source, read 2026-10-10. Workflow-versus-agent distinction, start-simple advice, and the trade-off between flexibility and cost or compounding errors.
- QS Digital checklist. QS's own method. The examples are hypothetical and describe no client.
