✕ Beranda Profil Langganan Per Project Proses FAQ Co-Researcher Blog Carousel Hubungi
Artikel ini juga tersedia dalam Bahasa Indonesia. Baca versi Indonesia →

Adding AI to a Live Website: Architecture Patterns Without a Rewrite

Adding AI to a Live Website: Architecture Patterns Without a Rewrite

When Your Manager Says "Let's Add AI"

Many developers are now getting the same request: a web application that has been running in production for years should suddenly "have AI features" — smarter search, automatic summaries, or a chatbot that understands internal data. The first temptation is to rewrite the system around the latest AI framework. Almost always, that is an expensive and risky decision.

The good news is that AI capabilities can be added incrementally. This article covers five architectural patterns that let an existing application gain AI features without a full rewrite: a sidecar service, asynchronous queues, pgvector in the database you already run, an MCP server for internal APIs, and feature flags plus quotas for a gradual rollout.

1. Sidecar Service: Keep AI Logic Out of the Main App

The first principle: never put calls to an AI model on the critical path of the main application. LLM calls are slow, their latency is unpredictable, and they can fail on provider rate limits. If that code lives inside the monolith, a single timeout can hold a web worker hostage and slow down the whole site.

The sidecar pattern moves AI logic into a separate service — say, a small Python or Node service — that the main application calls over internal HTTP. The benefits:

  • Failure isolation: if the AI service goes down, the main app keeps serving existing features and simply shows a fallback.
  • Separate dependencies: heavy AI libraries do not pollute the main application's dependency tree.
  • Independent scaling: you can add AI service instances without duplicating the web app.

The main application only needs to know one simple contract, such as POST /summarize with a tight timeout of a few seconds.

2. Queues and Async Jobs for Slow, Expensive Work

Not every AI task needs an instant answer. Summarising articles, generating images, or computing embeddings for thousands of documents should be pushed onto a queue. The user gets a fast response ("processing"), while workers do the heavy lifting in the background.

This pattern gives you several important controls:

Need How the queue helps
API cost Cap the number of workers to keep spending predictable
Provider rate limits Retry with backoff without burdening web requests
Reliability Failed jobs can be retried or moved to a dead-letter queue
User experience Pages stay responsive; results appear when ready

One caveat: a job that calls an external service is not always safe to retry automatically. If the operation has side effects (sending a message, for example), store a "processed" marker so nothing is duplicated.

3. Adding pgvector to Your Existing Database

Semantic search and RAG (Retrieval-Augmented Generation) need vector storage. You do not have to adopt a new vector database. If you already run PostgreSQL, the pgvector extension (Postgres 13 and newer) lets you store embeddings right next to your existing data:

CREATE EXTENSION vector;
ALTER TABLE articles ADD COLUMN embedding vector(1024);

The new column is nullable, so existing queries are unaffected. Backfill it gradually through queued jobs. For searching, pgvector provides distance operators such as <-> (L2), <#> (negative inner product), and <=> (cosine distance).

By default pgvector performs exact search with perfect recall. As data grows, add an HNSW index (better query performance, slower builds) or IVFFlat (faster builds). The pgvector docs recommend CREATE INDEX CONCURRENTLY in production so index creation does not block writes. Mind the dimension limits too: the vector type can be indexed up to 2,000 dimensions, while halfvec goes up to 4,000 at half precision.

Retrieval quality matters as well. Anthropic's research on Contextual Retrieval shows that prepending a short context (roughly 50–100 tokens) to each chunk before embedding cuts retrieval failures by 35%, by 49% when combined with contextual BM25, and by 67% once reranking is added. They also note that if your knowledge base is under 200,000 tokens, simply putting all of it in the prompt can be simpler than building RAG at all.

4. Wrap Internal APIs as an MCP Server

Instead of writing a bespoke integration for every assistant or model, consider the Model Context Protocol (MCP). MCP is an open protocol for giving LLMs secure access to tools and data sources. The modelcontextprotocol/servers repository holds reference implementations such as Fetch, Filesystem, Git, Memory, and Time, while official SDKs exist in many languages — including PHP, Python, TypeScript, Go, and Java.

That means the internal endpoints you already have (order status, product search, reports) can be wrapped as MCP tools. Any MCP-capable client can then use them with no extra integration work. Business logic stays in the existing API; the MCP server is just a thin layer on top.

The repository itself stresses that its servers are reference implementations, not production-ready solutions. Add authentication, prefer read-only tools where possible, and audit every call.

5. Feature Flags, a Rollback Plan, and Quotas

AI features should not ship to every user at once. Use a feature flag to enable them for admins first, then a small slice of users, and only then everyone. Prepare a clear rollback plan as well:

  • Turning off the flag must restore the old behaviour without a redeploy.
  • Database migrations stay additive (new columns, new tables), so nothing needs to be reverted.
  • Log every AI call — model, latency, tokens, status — for evaluation.

Finally, enforce quotas: per-user daily call limits and a monthly budget cap. Without them, one curious script or a looping bug can burn through your API budget overnight.

Conclusion

Adding AI to a live application is not about rewriting it; it is about adding layers with clear boundaries: AI logic in a sidecar, heavy work in a queue, embeddings in PostgreSQL via pgvector, data access through MCP, and everything guarded by feature flags and quotas. With these patterns, the existing system stays stable while new features are tested gradually and safely.

References

Share Article