✕ Beranda Profil Langganan Per Project Proses FAQ Co-Researcher Blog Carousel Hubungi
Artikel ini juga tersedia dalam Bahasa Indonesia. Baca versi Indonesia →

The Real Cost of AI Agents: Tokens, Servers, and Maintenance

The Real Cost of AI Agents: Tokens, Servers, and Maintenance

Why AI Agent Budgets So Often Miss the Mark

AI adoption in Indonesia is moving fast. An AWS study presented at AWS Summit Jakarta 2026 found that 40 percent of Indonesian companies now use AI in their operations, up from 25 percent the year before. Yet 56 percent of those companies are still in the early exploration stage, and measuring return on investment was named as one of the key challenges.

The root problem is often simple: budgets are built from subscription fees or the per-token prices on a pricing page, while an AI agent is a system that runs repeatedly and keeps reading context. The true cost only becomes visible after the agent has been live for a few weeks. This article breaks down each cost line so your AI project budget is grounded in reality.

The Main Cost Lines of an AI Agent

Broadly, there are four cost lines that must appear in your budget.

Cost Line What You Pay For Cost Behavior
Input/output tokens All text sent to the model (instructions, history, tool results) and all text the model generates Variable, grows with usage
Servers A VPS or cloud instance to run the agent gateway, or GPUs if you run local models Fixed monthly
Embedding & memory storage A vector database or memory files, logs, and skills Grows slowly with data
Monitoring Logs, alerting, token usage dashboards, security audits Fixed + team time

Tokens are the hardest line to predict. Input tokens usually far outnumber output tokens, because on every turn the agent resends its system instructions, tool definitions, conversation history, and tool results.

Servers are comparatively predictable. As a reference point, the documentation for OpenClaw, a popular open-source agent, lists a minimum of 2 GB RAM for the gateway alone with 4 GB or more recommended, plus 500 MB of disk at minimum and 2 GB or more for memory, skills, and logs. If you want to avoid API bills by running local models, requirements jump to GPUs: roughly 16 GB of VRAM for a 14B-class model that handles most tasks well, up to 40–80 GB (A100 class) for near-cloud quality.

Embedding storage is often treated as negligible, but every indexed document needs to be re-embedded whenever it changes. We will return to this in the maintenance section.

Monitoring is not optional. An agent with access to files, a shell, and business applications must be watched, both for cost and for security.

The Biggest Overlooked Cost Driver: The Heartbeat

Many business owners estimate cost based on the number of user requests. But an autonomous agent can burn tokens without a single user asking anything.

OpenClaw includes a heartbeat feature that fires every 30 minutes, around the clock, to check for pending tasks, and every cycle consumes tokens. Its own documentation identifies the heartbeat as the single biggest cost driver. Here are the cost estimates it publishes:

Usage Level Approx. Daily Cost Monthly
Light (CLI chat only) $1–5 $30–150
Moderate (heartbeat + channels) $5–20 $150–600
Heavy (many channels, complex skills) $20–50+ $600–1,500+
Local models (Ollama/vLLM) $0 in API fees $0 in API fees

The math is straightforward: once every 30 minutes means 48 cycles a day, or roughly 1,440 cycles a month. If each cycle carries 100,000 tokens of context, that adds up to 144 million input tokens per month just to check whether there is any work to do. With an isolated session mode that trims context to 2,000–5,000 tokens per cycle, the figure drops to roughly 3–7 million tokens.

Cost-reduction steps recommended in the OpenClaw documentation:

  • Use a cheap model (Haiku class) for the heartbeat, saving 80–90% of heartbeat cost.
  • Extend the interval to 60 minutes, saving 50%.
  • Enable quiet hours when there is no activity, saving 33%.
  • Run the heartbeat on a local model, bringing heartbeat API cost to zero.
  • Enable isolatedSession, cutting context from 100K to 2–5K tokens per cycle.

One documented case shows a bill falling from $1,200 per month to $36 through model routing alone. The budgeting lesson: ask your vendor or team how often the agent runs when nobody is using it.

A Bloated Context Window Means a Bloated Bill

Every token in the context window is paid for, on every turn. In its write-up on context engineering, Anthropic's Applied AI team describes context as a finite resource with diminishing returns. As token count grows, the model's ability to recall information accurately declines, a phenomenon known as context rot. So a bloated context is not only expensive; it can also degrade answer quality.

The guiding principle they offer is to find the smallest set of high-signal tokens that maximizes the likelihood of the desired outcome. In practice, that means:

  1. Concise but clear system instructions. Avoid stuffing in long lists of edge cases; a few diverse, canonical examples work better.
  2. Lean tool sets. Bloated tool sets with overlapping functionality are called out as one of the most common failure modes. Every tool definition also lands in the context.
  3. Just-in-time retrieval. Instead of loading every document up front, the agent keeps lightweight references (file paths, stored queries, links) and loads data only when needed.
  4. Compaction and tool result clearing. Long histories get summarized, and old tool results that are no longer needed are dropped from context.
  5. Structured note-taking. The agent writes progress to files outside the context window and reads them back when needed.
  6. Sub-agents. Sub-agents may explore using tens of thousands of tokens but return only a condensed summary of about 1,000–2,000 tokens to the lead agent.

Maintenance Costs That Rarely Make It Into the Budget

An AI agent is not a one-and-done project. There are recurring costs that should be budgeted as team hours:

  • Re-indexing embeddings. Whenever documents, product catalogs, or SOPs change, the data must be re-embedded. That carries embedding compute costs, plus the risk of a stale index that leads the agent to answer with outdated information.
  • Dependency and security updates. A real-world example from OpenClaw: 10 CVEs in 6 months, including a critical RCE vulnerability, and 341 malicious skills discovered in its community marketplace. The documentation advises always running the latest version, which means recurring time for patching, testing, and auditing.
  • Prompt and skill tuning. New models ship, behavior shifts, and business processes change. According to Anthropic, overly rigid prompts with hardcoded if-else logic increase maintenance complexity over time.
  • Reviewing agent output. Someone still has to check what the agent produces, especially in the early stages.

A practical rule: set aside explicit engineering hours per month in the budget, and never assume that number is zero.

How to Calculate an Honest ROI

A common mistake is comparing the API bill with one employee's salary. The honest comparison is total cost of ownership versus work hours that are genuinely replaced.

Total monthly cost = tokens + servers + storage + monitoring + (maintenance hours × hourly rate) + (setup cost ÷ useful life in months).

Value delivered = (task hours genuinely taken over by the agent − human review hours − error correction hours) × employee hourly rate.

An illustrative example with hypothetical figures:

Component Per Month
API tokens IDR 3,000,000
Servers + storage + monitoring IDR 1,000,000
Maintenance: 8 hours × IDR 150,000 IDR 1,200,000
Setup amortization (IDR 24M ÷ 12 months) IDR 2,000,000
Total cost IDR 7,200,000
Hours replaced: 80 hours − 20 review hours 60 hours
Value (60 hours × IDR 100,000) IDR 6,000,000

On paper the agent handles 80 hours of work, but once review time is subtracted, the value does not yet cover the cost. The answer is not to cancel the project outright, but to cut heartbeat costs, tighten the context, or widen the agent's scope until the numbers turn positive. Measure over at least 2–3 months, since the first month is usually dominated by setup and tuning costs.

Conclusion

The real cost of an AI agent rarely arrives as one big invoice. It comes from small things that repeat: a heartbeat every 30 minutes, a context window that keeps growing, and maintenance nobody budgeted for. By mapping the four main cost lines, controlling always-on agents, applying context engineering, and calculating ROI from the hours that are genuinely replaced, you can build an AI project budget that holds up.

Need reliable servers to run your AI agent or business website? The katili.dev team is ready to help, from hosting to development.

References

Share Article