8 min read
LLM integration services: add AI to an existing product without a rewrite
You rarely need a new stack to ship AI features. You need a clear integration boundary that handles permissions, structured output, cost limits and failure. Here is what that layer contains.
LLM integration services should add AI to your existing product through a thin, well-defined layer: a server-side endpoint that calls the model, validates structured output, enforces permissions and cost limits, and fails gracefully. You keep your stack, data model and auth. A rewrite is rarely needed; a clear integration boundary almost always is.
This guide describes what that integration layer contains and how to roll it out safely. It is based on implementation work, including adding an AI agent to the existing ImadDhin portal without changing its framework or data platform. The examples are code-level observations, not performance or revenue claims.
Why teams think they need a rewrite
AI features often start as a prototype outside the product: a notebook, a separate script, or an app built with an AI builder. It works, and the natural next thought is that the product must move to whatever stack the prototype used. Some vendors encourage this, because a new platform is a bigger project.
The other driver is fear of the existing codebase. Teams worry that their older backend cannot handle streaming, long requests or new infrastructure. In practice, most products already have what an LLM feature needs: authenticated users, an API layer, a database and a way to run background jobs. The model is just another external service, with some unusual properties.
What teams often get wrong
The most dangerous mistake is calling the model directly from the browser or mobile app with a provider key embedded in the client. Anyone can extract that key and run up your bill, and the client can send whatever context it likes.
The second is trusting model output. Text or JSON from a model is untrusted input. Writing it straight into the database, rendering it as HTML, or using it to choose which record to update invites errors and injection attacks.
The third is ignoring permissions during context assembly. If the feature pulls documents or records into the prompt, it must pull only what the current user is allowed to see. The model will happily summarize another customer's data if you hand it over.
- Provider keys in client code
- Unvalidated model output written to storage or rendered as markup
- Context assembled without the user's permissions
- One provider SDK imported across the whole codebase
- No timeouts, cost ceilings or per-user quotas
- Launching to every user at once with no evaluation set
A practical architecture
Put every model call behind one server-side module in your existing backend. The rest of the product talks to that module, never to a provider directly. A request through the layer follows the same steps every time.
- Authenticate the user and check they may use the feature
- Assemble context using only data the user is permitted to see
- Render a versioned prompt template with that context
- Call the model with a timeout and a bounded retry policy
- Validate structured output against a schema and reject what does not fit
- Apply guardrails and business rules before anything is saved or shown
- Record the request by reference, the prompt version, the model and the outcome
- Return the result, streaming where the user is waiting on text
Long-running tasks, such as processing a large document or running a multi-step agent, should become background jobs with a status the user can check and a notification when they finish. Put the whole feature behind a flag so you can release it to internal users, then a small group, then everyone, and switch it off without a deployment.
Implementation considerations
Choose features where AI output is easy to verify. Summaries, drafts, classification, field extraction and natural-language search are good first features because a user can see and correct the result. Features that take actions on the user's behalf need approval steps and belong later.
Use structured outputs where the provider supports them, and validate anyway. A schema turns free text into data your code can check. When validation fails, retry once with the error, then fall back to a clear message or a manual path rather than guessing.
Design the interface around review. Show AI output as a suggestion the user can edit, accept or discard, label it clearly, and keep the original input visible. Capture whether users accept, edit or reject results; that signal is the cheapest evaluation data you will ever get, and it shows which features are earning their place.
Plan for provider limits. Model APIs enforce rate limits and occasionally degrade. Honor retry-after signals, retry timeouts and server errors with bounded backoff, never retry validation or authorization errors, and decide in advance whether a fallback model is acceptable for each feature or whether the feature should simply pause.
Enforce quotas and cost limits on the server. Set per-user and per-feature limits, cap input size, and track token usage by feature. Implement counters atomically, with a transaction or an atomic increment, so concurrent requests cannot both pass the check. Cost is driven by input and output length, model tier and call volume, so measure those before choosing a model.
Review data handling. Check the model provider's terms on training use, retention and processing region, redact data the feature does not need, and document what leaves your system. Requirements depend on your sector and markets, so involve someone qualified to review them.
Build an evaluation set from real examples before launch: inputs, expected outputs and cases the feature should refuse. Run it on every prompt or model change. Model providers update and retire models, and a version change can shift behavior without any change in your code.
Write failure copy for people. When the provider times out or returns something unusable, tell the user what happened and what they can do next. A spinner that never ends, or a generic error with a stack trace, erodes trust in the feature.
Trade-offs
A provider abstraction adds a layer of indirection and can hide useful provider-specific features. Without it, switching providers or adding a fallback means touching every caller. A thin internal interface around the capabilities you actually use is usually the right compromise.
Hosted model APIs are fast to adopt and need no infrastructure; self-hosted open models give more control over data and cost at volume but add operational work and evaluation effort. Many products start hosted and revisit the decision only when volume or data requirements justify it.
Streaming improves perceived responsiveness for text generation but complicates validation, since you cannot fully validate structured output until it is complete. Stream prose, and validate structured results before acting on them.
Synchronous features feel immediate but tie up requests and fail visibly when a provider is slow. Background jobs are more resilient and easier to retry but need status tracking and notifications. Use synchronous calls for short, interactive tasks and jobs for anything that might take longer than a user will comfortably wait.
Lessons from ImadDhin work
These are implementation and code-level observations from the ImadDhin portal, not measured outcomes.
The portal's AI agent was added to an existing Next.js marketing site without changing its framework, hosting or database. Chat runs through the site's own API routes, and chat sessions, messages and saved memory are stored server-side with direct client access denied by security rules. The browser never talks to a model provider and cannot read or rewrite conversation history.
Different creative modes route to different model providers behind the same chat endpoint, and an older runtime path remains as a fallback. The model choice and assistant instructions come from remote configuration, so behavior can change without a deployment. Provider keys are resolved on the server from environment configuration or a secret manager, and none are exposed with a public prefix.
Premium features, such as live web research, are enforced on the API route rather than only hidden in the interface. A user who calls the endpoint directly without access is refused, which is the only kind of gating that matters.
When a provider fails, the agent shows user-facing error text instead of silently falling back to its introductory greeting. That small choice made failures visible during testing instead of looking like a strange but successful reply.
Common mistakes to test for
- Force a provider timeout and confirm the user sees a clear message
- Return malformed structured output and confirm nothing invalid is saved
- Include instructions in user content asking for other users' data
- Call the endpoint as a user without access to the feature
- Submit a very long input and confirm limits and costs are enforced
- Send parallel requests at the quota boundary and confirm the limit holds
- Switch the feature flag off and confirm the product still works
When a simpler solution is better
If a vendor you already use offers the AI feature you need inside their product, such as summarization in your helpdesk or drafting in your document tool, that may beat a custom integration. Your time is better spent on features specific to your product and data.
Some problems are not language problems. Better search indexing, clearer forms or a rules-based workflow can solve them more reliably than a model. And a nightly batch job that calls a model offline is often simpler than a real-time feature when users do not need results instantly.
Add AI without replacing your product
Pick one verifiable feature, build the integration layer properly once, and reuse it for the features that follow. If you want help adding AI to an existing product, explore ImadDhin's AI product development practice, send a project brief, or walk through your codebase and goals in a 30-minute call.
Frequently asked questions
Do we need to rewrite our product to add LLM features?
Rarely. Most products already have authentication, an API layer, a database and background jobs. A server-side integration layer that wraps model calls is usually enough to ship AI features on the existing stack.
Can we call the model directly from our mobile or web app?
Not with a provider key in the client. Keys can be extracted and abused, and the client can send any context it wants. Route calls through your backend, where you enforce authentication, permissions and limits.
How do we control LLM costs?
Enforce per-user and per-feature quotas on the server, cap input size, choose the smallest model that passes your evaluation set, and track token usage by feature. Make quota counters atomic so concurrent requests cannot bypass them.
What should an LLM integration project deliver?
A reusable server-side layer with permission-aware context, versioned prompts, output validation, logging, quotas and failure handling, at least one feature released behind a flag, and an evaluation set you can rerun on changes.
Should we self-host an open model instead of using an API?
Consider it when data control requirements or volume justify the operational work. Many teams start with a hosted API behind an internal interface, which keeps the option to switch later.
Ship AI features on the product you have
Walk through your product and the first AI feature worth building.
Book a 30-minute callIntegration layers, AI features and production rollout.
AI product developmentSix short steps; no account required.
Send a project briefKeep reading
AI product development in phases: from use case to production
AI products carry more uncertainty than conventional software. A phased approach, from use case to feasibility, narrow build, limited launch and expansion, turns that uncertainty into decisions backed by evidence.
Custom AI solutions vs off-the-shelf SaaS: a build-or-buy framework
Build where the workflow is a differentiator or the data is yours alone; buy where a vendor already solves a common problem well. Here is a five-question framework and the hybrid pattern most teams end up with.
AI agent evaluation before production: a practical evaluation harness
An agent that looked good in five demo conversations can still fail on the sixth real one. A small, repeatable evaluation harness turns quality from an impression into a report you can rerun on every change.