9 min read

MCP server development: when your product needs one and how to ship it safely

An MCP server turns your product's capabilities into tools that AI agents can call. Here is when that is worth building, when a simpler integration is enough, and how to ship one without handing agents more access than they need.

MCP server development makes sense when AI agents, yours or your customers', need a stable and permissioned way to call your product's capabilities. Build one when several agent clients need the same tools; skip it when a single internal integration will do. Treat every tool as a public API with scopes, typed inputs, and an audit trail.

The Model Context Protocol, or MCP, is an open standard for connecting AI applications to tools and data. A server exposes tools the model can call, resources it can read, and reusable prompts. Clients include desktop assistants, coding tools, and agent frameworks. The protocol makes integration easier; it does not make the tools behind it safe. That part is still your engineering work.

When your product needs an MCP server

An MCP server is worth building when the same capabilities must be reachable from several AI clients, and you want to define them once. Typical signals:

  • Customers ask to use your product from their AI assistant or coding tool, not only from your interface.
  • Your own agents run on more than one runtime or provider, and you want one tool layer shared across them.
  • A provider-managed agent session needs to call your systems, and MCP is the interface it accepts.
  • You want agent access to go through a documented, versioned, permissioned surface instead of ad hoc scripts against your internal API.

It is usually not worth it when one internal agent calls two functions in the same codebase. Plain function calling is simpler, and you can wrap those functions as MCP tools later if a second client appears.

Why MCP servers become risky

An MCP tool is an action a model can decide to take after reading content it did not write. The model may have just summarized an email, a product description, or a web page containing hidden instructions. If the tool behind that decision can reach far, a single manipulated input can do real damage within entirely legitimate permissions. The threat model is covered in more depth in securing agents.

Remote servers add ordinary API risks on top: authentication, token handling, rate limits, and logging. Community servers installed from a registry add supply-chain risk, because a server's tool descriptions and behavior can change after you trust it.

What teams often get wrong

  • Mirroring the whole API. Exposing every endpoint as a tool gives the model a huge, confusing surface and the widest possible blast radius. Agents work better with a few task-shaped tools.
  • One powerful token. A single service credential shared by every user means the server cannot tell who asked for an action, and any compromise reaches everything.
  • Relying on descriptions for safety. Writing "never delete records" in a tool description is a hint to a model, not a control. Enforce limits in code.
  • Returning raw data. Large, unfiltered responses waste context, leak fields the agent did not need, and can carry injected instructions back into the model.
  • Stateful sessions on stateless hosting. Holding per-session state in memory breaks on serverless platforms where each request may land on a new instance.
  • No idempotency on writes. Agents and clients retry. Without a key or ledger, a retried call can create a second order, message, or record.

A practical architecture for MCP server development

Keep the MCP layer thin. Your business logic, validation, and permission checks should live in a service layer that your web app, your API, and your MCP server all call. The MCP server translates protocol messages into those calls and back.

Design task-shaped tools

Start from what an agent needs to accomplish, not from your endpoint list. "Search products", "create a cart with these lines", and "list my recent orders" are better tools than generic create, read, update, and delete wrappers. Each tool should have a narrow input schema with limits on string length, array size, and allowed values, and a short, accurate description.

Separate reads from writes

Read tools can often run freely. Write tools need stricter treatment: allowlisted fields, value limits, an idempotency key, and for irreversible actions, a human approval step. Where possible, make writes staged rather than final, such as creating a cart and returning a checkout link instead of charging a card.

Authenticate per user, authorize per tool

For remote servers, the MCP authorization specification builds on OAuth, so users can grant a client scoped access to their own account. Behind the token, check on every call that the user is still entitled to the tool and that the target record belongs to them. For short-lived agent runs, a per-run credential that expires and is bound to one user and one run limits the damage of a leaked token.

Shape outputs

Return only the fields the agent needs, cap response size, and mark tool output clearly as data. Generic, user-safe error messages are better than raw provider errors, which can leak internal details.

Implementation considerations

  • Transport: local servers typically use standard input and output; remote servers use the streamable HTTP transport. On serverless hosting, a stateless server that returns complete JSON responses per request is the simplest reliable option.
  • Idempotency ledger: record each write call under a key derived from the run, tool, and arguments, created inside a transaction, so a repeated call returns the prior result instead of acting twice.
  • Rate limits and budgets: limit calls per user and per run, and alert on unusual patterns.
  • Audit logging: record who called which tool, when, the outcome, and a reference to any approval. Avoid logging secrets or full personal data.
  • Versioning: changing a tool's schema or meaning can silently break agents that learned the old behavior. Add new tools rather than changing existing ones in place, and announce removals.
  • Discovery: publish what the server offers, for example through a server card at a well-known URL, so clients and reviewers can see the surface without reading code.
  • Testing: exercise every tool with an MCP inspector and with scripted calls, including malformed input and repeated calls, before connecting a model.

Trade-offs

A public MCP server makes your product usable from many AI clients, and it also creates a permanent, externally used interface you must maintain and secure. A private server shared by your own agents gives consistency without that external commitment.

Stateless servers are easier to host and scale. Stateful sessions enable richer interactions such as server-initiated messages, at the cost of session storage and more complex hosting.

Strict idempotency protects writes and can be overly strict for reads. A ledger that treats a failed search like a failed payment will refuse harmless retries.

Lessons from ImadDhin work

The notes below are code-level observations from the MCP integration and public discovery files in this portal, not client outcomes.

  • Per-run credentials. The commerce agent reaches the portal through an MCP connection that accepts a random bearer token stored only as a hash, bound to one user and one run, and expiring soon after the run. Each request also checks that the run is still active and the user is still entitled to the agent.
  • A small surface. Only a handful of task-shaped commerce tools are registered on that connection, rather than a mirror of the whole API. Inputs are validated with typed schemas, including limits on quantities and the shape of product variant identifiers.
  • Staged, not final, writes. The cart tool returns a checkout link and never charges a customer, and a cart identifier is only accepted if it was created in the same account and run.
  • Transactional ledger. Each call is recorded under a key built from the run, tool, and arguments, created in a transaction. A completed call returns its stored result; an in-flight or failed one is reported as pending verification.
  • Reads and writes deserve different treatment. A common design mistake is a ledger that treats read tools the same as writes, so a failed search is refused on retry rather than simply re-run. Keep strict idempotency for writes such as carts, and let reads retry freely.
  • Visitor-safe discovery. The public site publishes a server card describing browser-based tools that only search sections, read the public services catalog, return contact details, or open pages; submitting an inquiry is limited to the Agent workspace.

Common mistakes to test for

  • Call a write tool twice with the same arguments and confirm only one side effect occurs.
  • Use an expired token, another user's token, and a token for a closed run, and confirm each is rejected.
  • Pass oversized strings, unexpected fields, and identifiers from another account.
  • Return a tool result containing instructions and confirm the agent treats it as data.
  • Revoke a user's entitlement mid-run and confirm the next tool call fails.
  • Confirm errors shown to the model and user contain no internal details or credentials.

When a simpler solution is better

If only your own application calls the capability, a normal API with typed function calling in your agent is simpler and easier to reason about. If a partner needs programmatic access but not through an AI client, a documented REST API with OAuth serves them better. Build an MCP server when AI clients are a real, recurring consumer of your product.

Ship a tool surface agents can use safely

Start with a few task-shaped tools, per-user authorization, staged writes, and a ledger, then widen the surface as real usage justifies it. If you are planning an MCP server for your product or store, see AI agent development and Shopify development, or talk it through in a 30-minute call.

Frequently asked questions

What is an MCP server?

It is a service that implements the Model Context Protocol to expose tools, resources, and prompts to AI clients such as assistants, coding tools, and agent frameworks, through a standard interface instead of custom integrations.

Does my SaaS product need an MCP server?

It helps when customers want to use your product from AI clients, or when several of your own agents need the same tools. If one internal agent calls a couple of functions, plain function calling is simpler.

How should a remote MCP server handle authentication?

Use per-user authorization based on OAuth, as the MCP authorization specification describes, then check entitlement and record ownership on every tool call. Avoid a single shared credential for all users.

How do I stop an agent from repeating a write through MCP?

Record each write under an idempotency key derived from the run, tool, and arguments inside a transaction, return the stored result for repeats, and flag uncertain outcomes for human verification rather than retrying automatically.

Is it safe to install community MCP servers?

Treat them like any third-party dependency with access to your data. Review what they can reach, pin versions, run them with minimal credentials, and watch for changes in tool descriptions or behavior.

Plan an MCP server you can defend

Review which capabilities should become tools and how each one is scoped.

Book a 30-minute call

Tool layers, MCP servers, per-run credentials, and audit trails.

See AI agent development

Commerce agents with catalog, cart, and order tools that never charge a card.

Explore Shopify development

Keep reading