A guide to the anatomy of effective commerce agents The architecture, latency & cost techniques, and eval practices for agents that make it easier to buy and sell online. Category Agents Product Claude Platform Date September 2, 2026 Reading time 5 min Share Copy link https://claude.com/blog/the-anatomy-of-effective-commerce-agents Author(s) Ali Shazal Matthew Koen Over the past year, we've worked with teams across the commerce industry — retailers, marketplaces, travel, entertainment, and telecom providers — to build commerce agents using Claude. These agents are in production, and enterprise customers have seen larger carts and more efficient seller operations when using them. They also share a simple architecture: Claude in an agent loop equipped with a set of skills, tools, and a strong eval suite. This post is for the engineers and engineering leaders building these (or other consumer facing) agents. Part 1 covers the architecture, which you decide once. Part 2 covers latency and cost. Part 3 covers production: memory, safety, evals, and scaling the work across an organization. Reference implementation We've also provided a blueprint to help build commerce agents on Claude. It contains the harnesses, patterns, and guardrails an engineering team needs to get a commerce agent running in days, with reference implementations of a shopping agent and a merchant agent for retail, travel, telecom, and ticketing platforms. anthropics/commerce-agents → In this guide Part 1: The architecture What is a commerce agent? Skills, not subagents System prompt or skill: decide by frequency Engineering agent tooling The UI components are tools Part 2: Making it fast and affordable Minimizing task completion latency Perceived latency Prompt caching Choosing the model and its configuration Part 3: Running it in production Memory that survives the session Safety: enforcement lives in the harness Evals: shipping a non-deterministic system Shipping with a large organization Looking ahead 01 The architecture One model in a standard agent loop, with skills for the long tail and tools that call the systems you already run. You decide this once. What is a commerce agent? We define a commerce agent as an agent that simplifies buying and selling across an online catalog. Some agents face consumers: they search, compare, substitute, and assemble the order. That could be a retail cart, a travel itinerary, a mobile plan change, or seats held for a show. Some agents face the business: they answer questions about sales, run promotions and campaigns, and manage inventory and pricing. The core architecture is a model in a standard agent loop : reasoning about a goal, exploring context, taking actions through tools, learning procedures through skills, asking clarifying questions, and observing the results until the goal is accomplished. There is no intent router in front of it that segments the conversation and no set of domain specific agents behind it. Engineering context Skills, not subagents A commerce agent has to cover a wide range of capabilities across many categories and intents, which makes it tempting to create one subagent per domain. In practice this proves suboptimal, because a commerce conversation is one tightly coupled session across multiple intents and turns, and requires considerable shared context. In a subagent architecture, the orchestrator holds the cart or staged changes, the user's preferences, and the conversation history. Every handoff to a subagent is a state-lossy operation, which often impacts the quality of the subagent’s response and, consequently, the overall response. On top of that, each handoff can cost several times the tokens and adds seconds of latency. The domains also rarely separate cleanly. A returns flow might need the order history, the current cart, and the product catalog, meaning a subagent-per-domain approach either duplicates that access everywhere or hands off mid-task. As models get smarter, they also handle longer context, more skills, and more tools, so the limits behind today's placement rules loosen with each model generation. Instead, agent skills give you similar per-domain modularity and context control without the handoff tax, because the skill instructions load into the main agent that already holds the entire history. In our comparisons across several enterprise deployments, a single agent with skills consistently has outperformed both the one-prompt-for-everything design and the subagent design on quality, and often at a lower cost and latency per task. Where subagents do earn their place is when the orchestrator can call them as a tool for a narrow or self-contained task that would benefit from its own dedicated context window. A common production example is a deep-research subagent, where the subagent searches and reads documents, writes and runs code, traverses data models, and hits dead ends. All the work happens inside one or more subagents, and only a compact answer comes back to the orchestrator. The other exception is a domain that already has its own purpose-built agent. If your pharmacy or financial-services experience runs a dedicated agent with its own compliance surface, the right move can be a hand-off, where that agent takes over the task and works with the user directly through its own loop until the task is done. The distinction is ownership of the conversation. A hand-off makes the domain agent the user's counterpart, while delegation keeps the orchestrator, bouncing the domain agent in and out within a single turn and degrading on every exchange. System prompt or skill: decide by frequency The main factor when deciding whether to put a set of instructions within a system prompt or skill is how often the agent will need it. Loading a skill costs a model turn, so anything the agent needs on most turns generally goes in the system prompt. This does, however, depend on how your traffic is distributed, and what agent behavior your evals show. A good starting point is that anything relevant to a third or more of your traffic, whether anticipated before launch or observed in production, goes in the system prompt, and the rest goes in skills. If a skill is predictable from a signal you already have, such as the page the user arrived from, we recommend injecting it from the harness before the first model call and skipping the extra turn to load the skill. Critical instructions, such as safety and legal rules, brand constraints, and key user facts such as allergies, always go in the system prompt. For commerce agents, this means product search lives in the prompt, since nearly every session touches it, and skills carry the long tail of features. In our reference implementation , the shopping agent's prompt holds grounding, cart and checkout semantics, and presentation rules, and the following skills cover the rest: search-discovery, purchase-research, planning-goals, customer-care, and memory-personalization. The merchant agent splits the same way, with performance-insights, catalog-listings, inventory-operations, pricing-promotions, and marketing-campaigns as its skills, one per operational domain. In the prompt Shopping agent Grounding, cart and checkout semantics, presentation rules, and product search. Shopping skills The long tail search-discovery · purchase-research · planning-goals · customer-care · memory-personalization Merchant skills One per operational domain performance-insights · catalog-listings · inventory-operations · pricing-promotions · marketing-campaigns Engineering agent tooling Our post on writing effective tools for agents covers tool design in general. Two points have mattered most in commerce: Build agent tools on top of your core systems and logic. A commerce company already has search and ranking, a cart, a preferences and profile store, an inventory system, promotion and campaign engines, sales analytics, and more, each encoding logic tuned over years and seeing signals the model never will. The agent's tools should call those systems, not reimplement them, and the tool boundary is where their logic ends and the model's judgment takes over. For example, when the agent calls search_products , the results should arrive already ranked; its job is to decide which results serve the user's goal, how many to show, and how to present them. Tool results are context. Return the fields the model reasons with and drop the rest. Image URLs on every search row are the usual offender. As needed, reshape the raw response inside the tool, including appending a next step when it isn't obvious from the data. This is especially relevant for error scenarios, where the model benefits from instructions instead of error codes. For example, add an error instruction "Include a product ID when querying availability," instead of a generic 403. The UI components are tools Most commerce agent responses are UI components rather than prose
A guide to the anatomy of effective commerce agents The architecture,
A guide to the anatomy of effective commerce agents The architecture, latency & cost techniques, and eval practices for agents that make it easier to buy and sell online. Category Agents Product Claude Platform Date September 2, 2026 Reading time 5 min Share C
这条信息对 FDE 的直接价值在于提醒交付人员持续关注模型、智能体与企业流程之间的变化。面对类似项目,应先确认客户的真实业务目标、数据边界、权限条件和验收指标,再选择工具并用最小场景验证结果,避免只追逐功能更新。