9 min read
Conversational interface framing is too small to describe agentic products.
Because the quality of agentic products entirely depends on the infrastructure around that interface.
The LLM is often being asked to operate inside a half-built product with no durable memory, no recovery path, no tool boundary, no cost discipline, and no idea what happened five steps ago.
The real product is a loop:
This loop needs infrastructure and your agentic stack should allow you to build it with a set of layers that make that loop reliable enough to deliver at every user interaction.
Let's have a look at 7 layers to build truly great products.
Finally, I will share the default stack I would start with and present options for scale.
Let's dive in.
We cover a lot more in our Agentic SaaS Playbook, and we go further inside The Agent Foundry on framework comparisons, model routing, memory architecture, ingestion pipelines, MCP tool safety, observability, deployment, and implementation checklists.
Traditional SaaS stacks became boring in a good way.
Postgres, React, object storage, queues, auth, billing, observability. You can still make bad choices, but the shape is stable.
Agentic stack is less settled because the model is an active participant in the control flow.
That changes the stack and now you need layers for:
The mistake that I see often is treating these as optional add-ons whereas they are the actual product boundaries.
If your agent cannot remember, retrieve, call tools safely, resume after failure, explain what it did, and keep costs inside a margin envelope, you just have a demo that invoices your API key.
The first stack decision is "what owns the loop?"
An agent framework gives you the control structure for planning, tool calls, state, retries, handoffs, and intermediate steps.
You can write that loop yourself, but most teams should not start there.
Hand-rolled loops are fun until you need human approval, partial retries, tool-call validation, branch recovery, memory injection, and traces across 14 steps.

For example, a practical choice may look like this:
The important part is picking one execution model and staying consistent, because every framework has its own idea of state, tools, memory, traces, and recovery.
Model selection is now cost architecture.

The spread between cheap routing models and frontier reasoning models is enormous (e.g., Opus vs DeepSeek).
If every request goes to the strongest model, your margins are worse than they need to be.
The default production pattern should be cascade routing:
cheap model -> standard model -> frontier model
Use the cheapest model that can plausibly do the job, then escalate when the task is complex, high-risk, or low-confidence.
A simple version looks like this:
type TaskTier = "route" | "standard" | "frontier";
function chooseModel(input: {
taskType: string;
risk: "low" | "medium" | "high";
requiresLongContext: boolean;
userPlan: "free" | "pro" | "enterprise";
}): TaskTier {
if (input.risk === "high") return "frontier";
if (input.requiresLongContext) return "standard";
if (input.taskType === "classify" || input.taskType === "extract") return "route";
if (input.userPlan === "enterprise") return "standard";
return "route";
}That is not a perfect router but a good first router because it forces the product to admit that not all requests deserve the same model.
Then add caching:

The order matters here and routing and caching should arrive before fine-tuning. They are easier to ship, easier to measure, and usually more important for gross margin.
Once you get familiar with your task space and become comfortable with the idea of model routing, you can try more advanced alternatives such as vLLM Semantic Router v0.2 Athena.
Agents need memory because customers expect continuity.
They expect the product to remember their company, their data, their preferences, their workflows, their last run, and the exception they already explained three times.
Basic RAG gives you document recall which is useful, but it is not the whole memory layer.
A real memory system has several kinds of state:
For most early products, Postgres plus pgvector is the right default.
It keeps relational data, metadata, permissions, and embeddings in one operational system, and you also avoid dual-write bugs between your app database and a separate vector database.
Move to a dedicated vector system when you have a real scale or retrieval-quality reason, not because your architecture diagram wants another logo.

The retrieval pattern should also mature:
query -> rewrite -> retrieve -> rerank -> compress -> answer -> cite
The strong version decides what to search, reranks the candidates, compresses context, cites sources, and records what was useful.
As models grow their context size, you can also convert some of your workloads to leverage long-context instead of building a RAG layer.
Once you reach the RAG or long-context milestone, you will find yourself building a memory layer, which is entirely another topic on its own.
Conversational memory has evolved far beyond appending the last N messages to a prompt.
Modern agent memory systems provide long-term storage that persists across sessions, with semantic retrieval and structured knowledge extraction.

For most SaaS products, the pragmatic memory architecture combines three layers: (1) a conversation buffer for the current session (last N messages), (2) a vector memory store for long-term retrieval of facts and documents, and (3) a structured knowledge graph for entity relationships and temporal queries. Start with layers 1 and 2; add layer 3 only when you see the bot "forgetting" important details across sessions.
Bad ingestion creates bad agents.
It does not matter how good your model is if the PDF parser drops tables, the CSV pipeline loses column semantics, the website crawler captures navigation junk, or the sync job embeds stale data without lineage.
Ingestion should answer five questions:
That means the ingestion stack needs more than "upload file, split chunks, embed."
You need connectors, parsers, chunking strategy, PII handling, lineage, retries, and incremental sync.

The practical default:
Most agent quality problems that look like "the model hallucinated" are actually ingestion problems wearing a model costume.
Tool use is where agents leave the text world and touch reality.
That makes it the highest-risk part of the stack.
A tool is not just a function. It is a contract:
The unsafe pattern is letting the model improvise against loosely described tools.
The safer pattern is a tool gateway:
agent -> tool policy -> schema validation -> execution sandbox -> audit log -> result validation
Use MCP where possible, but do not confuse protocol support with safety. MCP makes tools portable. It does not automatically make them safe.

For production SaaS, tools need policy.
The more valuable the workflow, the more explicit the tool boundary should be.
Agents fail in expensive ways.
A normal API request fails and returns a 500.
An agent run can fail after 12 model calls, 4 retrieval passes, 3 tool calls, and one half-finished external action.
If the only recovery strategy is "start again," you are burning money and creating duplicate side effects.
Production agent runs need durable execution:
Framework-level persistence helps, but long-running product workflows often need a durable execution engine too.
Temporal, Inngest, Step Functions, and similar systems matter because they turn agent work into recoverable jobs instead of fragile request lifetimes. ReLiveGym extends that recoverability question to timing itself, testing wake-up trigger mechanisms that decide when a long-lived agent should resume work.

The rule is simple:
If a run can outlive a request, cost real money, or touch an external system, it needs durable execution.
You cannot debug agents with plain logs.
You need to see the run as a structured object:
Without this, every production issue becomes a séance.

The useful primitive is a run_id that follows the work everywhere: logs, traces, costs, tool calls, evals, and customer-visible history.
{
"run_id": "run_01J...",
"tenant_id": "tenant_123",
"workflow": "invoice_reconciliation",
"model": "standard",
"retrieval_chunks": 8,
"tool_calls": 3,
"cost_usd": 0.042,
"status": "completed"
}Then add product evals:
Observability tells you what happened, and evals tell you whether it was good.
You need both.
For a serious but still small team, I would keep the first stack boring:

The goal is to preserve optionality while you are still learning what the product actually is.
Do not start with a platform you cannot operate, or self-host frontier models before you understand utilization.
Adding multi-agent choreography before a single agent can complete the job reliably will burn you out.
Forget about the second database unless retrieval quality or scale forces you.
The winning products are the ones where the agent loop is wrapped in enough boring infrastructure that customers can trust it with real work.
For the complete version, read the full Agentic SaaS Stack chapter inside The Agent Foundry.