Introduction
Every team that starts building an LLM-powered product eventually hits the same question: where should this extra knowledge live? Is memory enough, do we need a RAG pipeline, or should we fine-tune the model? The three are often treated as interchangeable, but each one solves a different problem.
The most useful mental model comes from what Anthropic calls context engineering. Language models have a limited attention budget: as more tokens are packed into the context window, the model's ability to accurately recall information from it declines, a phenomenon known as context rot. So the real challenge is not simply adding knowledge, but getting the most relevant tokens into context at the right moment, at the lowest possible cost. Memory, RAG, and fine-tuning are three different answers to that challenge.
1. Memory: Long-Lived Facts and Preferences
Memory is a small, curated set of notes carried across sessions. It is not a document store. It holds compact facts that are almost always relevant: user preferences, project conventions, environment details, or lessons learned from past mistakes.
Hermes Agent by Nous Research offers a concrete example. Its memory consists of just two files: MEMORY.md for the agent's own notes (capped at roughly 2,200 characters, or about 800 tokens) and USER.md for the user profile (roughly 1,375 characters, or about 500 tokens). Both are injected into the system prompt as a frozen snapshot at the start of each session, which keeps the prefix cache intact. The agent manages them itself through a tool with add, replace, and remove actions.
Those size limits are intentional. The Hermes documentation advises against storing facts that are easy to look up again, raw data dumps such as logs or tables, and ephemeral details that only matter within a single session. For long conversation history, Hermes relies on a separate mechanism: session_search, backed by SQLite FTS5, which is queried only when needed. Because memory lands in the system prompt, new entries are also scanned for prompt-injection patterns before they are accepted.
Anthropic describes a similar pattern as structured note-taking, or agentic memory: the agent writes notes outside the context window, such as a NOTES.md file, and pulls them back in later. The Claude Developer Platform also offers a file-based memory tool for maintaining project state across sessions.
Use memory when: the information is small, stable, tied to a specific user or project, and needed in nearly every interaction.
2. RAG: Large, Frequently Changing Knowledge Bases
Retrieval-Augmented Generation (RAG) fetches relevant pieces of information from a knowledge base and appends them to the prompt. The standard pipeline splits documents into chunks of a few hundred tokens, converts them into embeddings, stores them in a vector database, and retrieves them by semantic similarity at query time.
A few key lessons from Anthropic's Contextual Retrieval research:
- Check whether you need RAG at all. If your knowledge base is under 200,000 tokens (about 500 pages), you can include all of it directly in the prompt. With prompt caching, latency can drop by more than 2x and costs by up to 90%.
- Combine embeddings with BM25. Embeddings are great at capturing meaning but can miss exact matches like the error code
TS-999. BM25 handles that kind of lexical lookup. - Add context to every chunk. A chunk like 'revenue grew by 3% over the previous quarter' is useless without knowing which company and period it refers to. Contextual Retrieval prepends a short explanation of roughly 50 to 100 tokens to each chunk before indexing, cutting retrieval failures by 49%. Adding a reranking step pushes that reduction to 67%.
RAG's main strength is freshness. When prices, documentation, or policies change, you update the index without touching the model. Answers can also be traced back to their source documents, which matters a great deal for auditability and user trust.
Use RAG when: the data volume is large, it changes often, and answers must be accurate and verifiable.
3. Fine-Tuning: Changing Style and Behavior, Not Adding Facts
Fine-tuning updates the model's weights using new training examples. Its impact is most visible in how the model responds: consistent output formats, brand voice, classification behavior, or reasoning patterns for narrow, repetitive tasks.
The most common mistake is using fine-tuning to inject factual knowledge, such as hosting plan prices or product documentation. That approach causes problems: facts baked into weights cannot be updated without retraining, the model cannot show where an answer came from, and it can still confidently invent plausible-sounding details. For facts, RAG is almost always cheaper, easier to keep current, and safer.
Before committing to fine-tuning, make sure the lighter options have been exhausted. Anthropic recommends starting with a minimal prompt on the best available model, then adding clear instructions and a curated set of diverse, canonical examples based on the failure modes you observe during testing. Hermes Agent also shows that an agent's default personality and voice can be shaped through a SOUL.md file without any retraining at all.
Use fine-tuning when: prompting and examples are no longer enough, you need consistent behavior at scale, or you want to distill a capability into a smaller model to cut cost and latency.
4. Decision Table
| Need | Data Change Frequency | Volume | Accuracy Requirement | Approach |
|---|---|---|---|---|
| User preferences, project conventions | Rarely, grows slowly | Very small (hundreds of tokens) | Consistent across sessions | Memory |
| Small reference documents | Moderate | Under ~200K tokens | High | Full context + prompt caching |
| Docs, FAQs, policies, catalogs | Frequently (daily/weekly) | Large and growing | High, needs sources | RAG (embeddings + BM25 + reranking) |
| Error codes, SKUs, invoice numbers | Frequently | Large | Must match exactly | RAG with BM25 / contextual BM25 |
| Past conversation history | Constantly growing | Unbounded | Recall of specific details | Session search / retrieval over history |
| Tone, output format, writing style | Almost never | Hundreds to thousands of examples | Behavioral consistency | Prompt + few-shot, then fine-tuning |
The rule of thumb is simple: the more often data changes, the further it should live from the model's weights. Data that changes daily belongs in RAG, stable personal context belongs in memory, and only behavior that almost never changes is worth baking in through fine-tuning.
5. The Most Common Production Combination
In production, teams rarely pick just one. The most common setup is a solid system prompt + RAG + memory, with fine-tuning as the last resort. Here is the implementation order we recommend:
- Baseline prompt and evals. Write a clear system prompt, add a few examples, and build a test question set. Without evals, you cannot tell whether the next step actually improves anything.
- Full context + prompt caching while the knowledge base is still small. This follows Anthropic's advice to do the simplest thing that works.
- RAG once data becomes too large or changes too often. Start with hybrid embeddings + BM25, then add contextual retrieval and reranking if accuracy is still falling short.
- Memory for personalization and cross-session continuity, with strict size limits and clear control over what may be saved.
- Fine-tuning only when evals reveal behavioral or formatting issues that prompting cannot fix.
For example, a support chatbot for a hosting service might use RAG for documentation, tutorials, and plan pricing; memory to remember that a particular customer runs Laravel on a specific PHP version; and the system prompt to keep the tone on-brand. Fine-tuning only enters the picture if, say, you want to move ticket classification to a smaller, cheaper model.
Conclusion
Memory, RAG, and fine-tuning are not competitors. They are tools for different layers. Memory answers who the user is and what their working context looks like, RAG answers what the facts are right now, and fine-tuning answers how the model should behave. Start with the simplest option, measure with evals, and add complexity only when the data proves you need it.
Need help building a website or application that is ready for a RAG-powered chatbot? The katili.dev team can help, from hosting infrastructure to implementation.