✕ Beranda Profil Langganan Per Project Proses FAQ Co-Researcher Blog Carousel Hubungi
Artikel ini juga tersedia dalam Bahasa Indonesia. Baca versi Indonesia →

Self-Host LLM vs API: An Honest Cost Breakdown for Indonesian Businesses

Self-Host LLM vs API: An Honest Cost Breakdown for Indonesian Businesses

The Question We Keep Hearing

Over the past few months, the questions from our clients have shifted. It used to be: should my business use AI at all? Now it is: which is cheaper, running a large language model (LLM) on our own servers or continuing to pay for an API from providers like OpenAI, Anthropic, or Google?

The honest answer: it depends on your volume, the kind of data you handle, and how ready your team is to operate infrastructure. This article is not here to sell you either option. We will break down each cost component, calculate the break-even point with a formula you can reuse, and then walk through the hybrid pattern we most often recommend for small and mid-sized businesses.

Note: every price in this article is an illustrative figure meant to make the math easy to follow. GPU prices, API rates, and exchange rates move quickly, so always check current pricing before you commit.

The Real Cost Components

API Costs: Paying per Token

With an API, you pay per token (a chunk of a word) going into and coming out of the model. There are no GPU servers to manage, but your bill grows linearly with usage. The community documentation for OpenClaw, an open-source AI agent that can connect to many model providers, gives a realistic picture: light usage can run around USD 30–150 per month, moderate usage USD 150–600, and heavy usage can exceed USD 1,500 per month. Interestingly, the biggest cost driver is background automation that runs around the clock, not manual chatting.

Self-Hosting Costs: GPUs, VPS, Power, and Bandwidth

Self-hosting means running an open-weight model (such as Qwen or Llama) on your own infrastructure using a runtime like Ollama, vLLM, or SGLang. The model may be free, but the hardware is not. The same guide maps rough VRAM requirements: 8 GB for an 8B model (basic tasks), 16 GB for 14B, 24 GB on an RTX 4090-class card for 32B, and 40–80 GB on A100-class hardware for 70B models that approach cloud quality.

A self-hosted setup typically includes:

Component Cloud GPU / GPU VPS Your Own Server (On-Premise)
Hardware Rented hourly or monthly, no upfront capital Bought upfront (capex), depreciated over 3–4 years
Electricity Included in the rental price Roughly IDR 1,400–1,700/kWh depending on PLN tariff, plus cooling
Bandwidth Usually a bundled quota, egress may be billed Needs a dedicated line and a stable public IP
Application server Separate VPS for the backend Can be combined, but riskier
Labor Setup, updates, monitoring Same, plus physical maintenance

A rough on-premise power example: a machine with a 4090-class GPU draws about 0.6 kW under load. Running 24/7, that is around 430 kWh per month, or roughly IDR 600,000–730,000, before counting air conditioning for the server room.

For the application side (backend, database, model gateway), a cloud VPS remains a sensible foundation. ScalaHosting points out that a cloud VPS provides dedicated resources unaffected by other tenants and can be scaled in seconds, with solid plans typically starting around USD 20 per month.

Break-Even: When Does Self-Hosting Become Cheaper?

The simple formula:

Break-even volume (tokens/month) = (Infrastructure cost + Operational cost) ÷ API price per token

Suppose a data-center-class cloud GPU runs 24/7 for about USD 1,100 per month, plus operational overhead (part of an engineer's time for updates and monitoring) of around USD 700. That totals USD 1,800 per month.

Comparable API price (per 1M tokens, blended) Monthly break-even volume Notes
USD 0.50 (lightweight model) 3.6 billion tokens Likely beyond one GPU's capacity, so the API almost always wins
USD 3 (mid-tier model) 600 million tokens Realistic for businesses with high-volume support chatbots
USD 10 (premium model) 180 million tokens But open-weight quality may not be equivalent

For a sense of scale: a typical customer service conversation might consume about 2,000 tokens. Ten thousand conversations a day adds up to roughly 600 million tokens a month, right at the break-even point of the mid-tier scenario.

There is an important trap here: compare like with like. A 32B model you host yourself should be compared against an API model of similar capability, not against the most expensive frontier model. Many providers now sell hosted access to open-weight models at very low per-token prices, and at that point pure self-hosting rarely wins on cost alone.

Non-Cost Reasons: PDP Law, Sensitive Data, and Model Control

Often, the decision to self-host is not about saving money but about compliance and control.

Indonesia's Personal Data Protection Law (Law No. 27 of 2022). The law regulates transfers of personal data outside Indonesia, which in principle require that the destination country offers equal or higher protection, or that adequate and binding safeguards are in place, or that the data subject has consented. When you send customer data to an API hosted abroad, you are effectively making a cross-border transfer. That does not make APIs off-limits, but you need a legal basis, a data processing agreement, and proper documentation. Check the specifics with your legal counsel.

Sensitive data. Medical records, financial data, legal documents, and trade secrets are strong candidates for processing on your own infrastructure. The OpenClaw documentation lists privacy as a core reason for its local-first design: with a local model, data does not have to leave your machine.

Full control over the model. API providers can change model behavior, retire older versions, or raise prices. When you self-host, you pin the model version, can fine-tune it, and are not subject to a third party's policy changes.

The Hidden Burden of Self-Hosting

This is the part that planning spreadsheets most often underestimate.

  1. Updates and security patches. The open-source AI ecosystem moves extremely fast. For example, the OpenClaw community docs record 10 CVEs in six months, including a critical remote code execution flaw, and more than 40,000 gateways found exposed on the internet without adequate protection. The lesson applies broadly: anyone who self-hosts must update routinely, restrict ports, and enable authentication.
  2. Monitoring. You need to watch GPU utilization, latency, request queues, and errors. The official OpenClaw documentation even ships dedicated commands for gateway status checks, doctor, and triage, because day-to-day operations genuinely require diagnostic tooling.
  3. Capacity planning. Traffic is uneven. If business-hour peaks need three times your average capacity, your GPU sits idle overnight while still being billed.
  4. On-call duty. When the model server goes down at 2 a.m. and takes your customer chatbot with it, who gets woken up? With an API, that is the provider's problem.
  5. Model lifecycle. New open-weight models ship every few months. Each upgrade means re-evaluating quality, running regression tests, and possibly meeting different VRAM requirements.

A hosting analogy fits well here. ScalaHosting explains that an unmanaged VPS gives you complete freedom along with complete responsibility when something breaks, while a managed VPS hands the technical upkeep to the provider. Self-hosting an LLM is the unmanaged version of AI.

The Hybrid Pattern We Recommend Most Often

For most small and mid-sized businesses, the best answer is not one or the other, but a combination.

Layer Role Runs on
Gateway / model router Request routing, failover, token usage logging Your own cloud VPS
Small local model (8B–14B) Personal data masking, classification, internal summaries, routine tasks GPU VPS or local server
Frontier model API Complex reasoning, public-facing content, high-quality tasks API provider

How it works in practice:

  • Sensitive data stays in. The local model masks names, national ID numbers, phone numbers, and other personal data before any request is forwarded to the API.
  • Routine work goes to cheaper models. The OpenClaw documentation cites one case where costs dropped from USD 1,200 to USD 36 per month through model routing alone, sending lightweight tasks to cheaper or local models.
  • Two-way failover. If the API has an outage, critical tasks can fall back to the local model; if the local server is down, some tasks can be shifted to the API.
  • Quarterly reviews. Track monthly token volume. Once you approach break-even, move your largest and most stable workloads to self-hosting.

For small businesses, start with an API plus a gateway on a small cloud VPS, and add a local model only when you have a genuine sensitive-data need. For mid-sized businesses processing hundreds of millions of tokens per month, a single GPU node for routine tasks combined with an API for heavy lifting usually strikes the best balance between cost, compliance, and team workload.

Conclusion

Self-hosting an LLM is not automatically cheaper, and an API is not automatically easier in the long run. Measure your token volume, compare against models of similar capability, include labor in the formula, and factor in your PDP Law obligations. For most Indonesian businesses today, a hybrid pattern offers the flexibility to save money without compromising data security. If you need help setting up a cloud VPS, a model gateway, or a GPU server, the katili.dev team is ready to design it around the scale of your business.

References

Share Article