The short answer
If you run an MCP server for your own product and traffic is light, your bill usually comes down to three things: the box it lives on, the LLM tokens flowing through it, and your own time. Self-hosting keeps the box cheap but bills you in evenings and weekends. A managed MCP host charges more per month but trades your attention for uptime, dashboards and per-client caps. Neither is “free.” The trick is matching the option to the stage your product is actually in.
Below is a grounded comparison for an indie founder or a small team: what each path costs in money, what it costs in time, and where token spend quietly eats margins.
What you are actually paying for
An MCP server itself — the open-source piece that brokers between your AI clients and your tools — does not have a license fee. The reference implementations and most community servers are MIT or Apache licensed, so the software line item is essentially zero.
The real costs are operational and downstream:
- A host to run the server on (a small VPS, a container service, or a managed MCP provider’s plan).
- LLM token usage when your agent or your users’ agents call tools and the model reasons about the response.
- Engineering hours for patching, OAuth refresh flows, version bumps of the upstream MCP SDK, and watching the thing at 2 a.m.
- Third-party API costs that flow through the MCP server (Stripe, Notion, your database, your vector DB).
Industry estimates for fully custom MCP servers built against specific SaaS integrations often land in the $50,000–$150,000 per integration per year range once you fold in development, QA, monitoring and ongoing maintenance. For an indie founder, you will not spend that much, but the shape of the cost — operations and maintenance dominate — still holds.
Self-hosting at indie scale
Self-hosting means you take the open-source MCP server, point it at the tools and APIs you want to expose, and run it on infrastructure you control. Usually that is a $5–$20/month VPS from Hetzner, DigitalOcean or a comparable provider, plus the model API keys you already pay for.
The direct cost picture:
- A small VPS capable of running an MCP server and a reverse proxy is usually in the $5–$20/month range.
- Bandwidth is modest for low traffic; at hobby scale you will not notice it.
- Your LLM bill is whatever your model provider charges for tokens in and out. Token spend is not a function of MCP itself — it is a function of how chatty your prompts and tool definitions are.
The hidden cost is time. Industry analyses of self-hosted platforms suggest operations and maintenance routinely represent more than half of the total cost of ownership, and that security patching alone can consume hundreds of developer hours a year per team. For a solo founder, even a small fraction of that — say a few hours a month of dependency updates, OAuth debugging, and chasing the occasional broken upstream — is the real price of self-hosting.
The advantages are real too. You own the runtime, you can audit every tool call, you can put the server behind your VPN, and you are not on the hook if a managed provider changes pricing.
The risks:
- No one wakes you up when it crashes. You need at least basic uptime monitoring.
- OAuth refresh, credential rotation, and upstream SDK churn land on your plate.
- If you ever serve real customers, the implicit SLA is “whenever the founder is awake.”
A managed MCP host at indie scale
Managed MCP hosting means a provider runs the server lifecycle for you. You configure the tools, plug in credentials, and connect clients like Claude, Cursor or ChatGPT over a remote endpoint. The provider handles uptime, scaling, patches, and usually gives you a dashboard.
What it typically costs:
- Workspace plans that start at $0 for a free tier with a small monthly execution allowance.
- Paid plans in the tens of dollars per month that lift execution caps, add SSO and basic observability.
- Enterprise tiers priced per seat or per usage once you need SSO, audit logs and procurement paperwork.
- Some providers also offer scoped engagement pricing — fixed-fee projects that build and launch an MCP endpoint on a branded domain — often starting in the low thousands of dollars and quoted after a scoping call.
What you are buying is not the server. It is the operator. Patching, scaling, basic rate limiting, credential storage and the on-call rotation belong to someone else.
The trade-offs:
- Your data transits a third party’s infrastructure. Read the provider’s data handling terms before you send anything sensitive through.
- You depend on that provider’s uptime. If they go down, your agents go down.
- Per-seat or per-execution pricing compounds as you grow. The plan that felt cheap at three users can sting at thirty.
- Some providers do not offer a self-hosted edition, which can lock you in if you ever need to migrate.
Where token spend actually goes
Token cost is the line founders most often mis-model. MCP itself does not add a token tax on top of the model — what it does is make it easy to call many tools in one conversation. Each tool call sends the tool description to the model, the model reasons about it, and the response comes back through the same pipe. Three patterns drive most of the spend:
- Tool description bloat. Every tool you expose adds its schema to the prompt. If you have fifty tools, you are paying for fifty schemas on every turn. Curate the surface area.
- Verbose tool outputs. If a tool returns a full document or a large JSON blob and the model echoes part of it back, you pay twice. Trim what the tool returns to what the agent actually needs.
- Idle polling or always-on workflows. An agent that calls a tool every minute to “check” something will burn tokens whether or not anything changed. Batch checks or use event-driven triggers instead.
The practical levers for indie founders:
- Per-client caps: set a daily or monthly token budget per connected client, so a runaway agent cannot drain the bill.
- Idle shutdown: turn the MCP server (or the LLM call path) off when no clients are connected, or put it behind a cron-style schedule.
- Cheaper models for tool selection: use a small, fast model to pick which tool to call and a larger model only for the heavy reasoning steps.
- Tool-result truncation: cap the size of any tool response that re-enters the model.
These work the same way whether you self-host or pay a provider. The difference is who configures them — you, or your provider’s dashboard.
How to think about idle shutdown
Idle shutdown is one of the cheapest cost controls available to a low-traffic MCP deployment, and it is the one founders skip most often.
- On a VPS: a simple health check can stop the container when there have been no requests for N minutes and start it again on the next connection. Cold-start latency is the cost; a few seconds of delay is the price of not paying for an idle process.
- In a managed setup: many providers expose sleep or scale-to-zero behavior on their lower tiers. If yours does not, ask whether the plan bills for provisioned time or for actual executions — that single detail decides whether idle shutdown is worth scripting on your side.
If your MCP server only serves one user at predictable times (your own dev work, say 9–6 weekdays), turning it off the rest of the week can cut a meaningful slice of your hosting or compute bill with no real downside.
Self-host versus managed: a decision frame
Use this frame the next time you are tempted to flip the comparison into a binary:
Pick self-hosting when:
- You have one to a few trusted users.
- Traffic is low and bursty, not customer-facing.
- You have the time and inclination to patch, monitor and debug.
- The data touching the MCP server is regulated or proprietary enough that you need it on your own network.
- You want to keep unit cost flat as usage grows.
Pick a managed host when:
- The MCP server is part of a product you sell and uptime actually matters.
- You would rather pay $20–$100/month than learn OAuth refresh flows.
- You need SSO, audit logs, or per-tenant caps out of the box.
- You want to spend your evenings shipping the product, not babysitting a container.
A hybrid is also reasonable early on: a managed host for production, a self-hosted copy for development and testing. The interesting architectural question is whether the deployment model is a configuration choice — something you can flip — or a structural one baked into your code. If it is a configuration choice, you can change your mind later without rewriting the integration.
A 30-minute setup that keeps costs visible
If you want a concrete next step, here is a small loop that works for either path:
- Pick one model provider and one MCP server image. Wire them together with a single tool that reads from one API you own.
- Add a per-client daily token budget and a hard stop when the budget is hit.
- Add an idle shutdown or scale-to-zero rule with a five-minute timeout.
- Log every tool call with input size, output size, model used, and total tokens. Even a CSV is fine at this stage.
- After a week, look at the log. You will see where the tokens actually go, and which tool descriptions are doing nothing.
That loop will tell you, in numbers, whether the operational savings of a managed plan outweigh its monthly fee at your current traffic. It is also the kind of evidence you can show a customer or an investor when they ask what your AI infrastructure actually costs.
FAQ
Is a managed MCP host just a wrapper around the same open-source server? Usually yes, plus operational tooling — auth, dashboards, scaling, audit logs. You are paying for the operator, not for new server code.
Do I need to pay anything to run an MCP server at all? The server software is open source and free. You pay for the box it runs on, the model tokens it consumes, and any third-party APIs it calls. A free tier on a managed provider can cover very light personal use.
Can I migrate from managed to self-hosted later? Sometimes. A few providers offer a self-hosted edition of their product; many do not. If portability matters, treat the deployment model as a configuration choice from the start, and confirm before you commit.
What is the single biggest cost control to implement first? Per-client token caps with hard stops. They are five lines of config and they protect you from the most common failure mode — a runaway agent that calls tools all night.
Sources
- https://truto.one/blog/build-vs-buy-the-hidden-costs-of-custom-mcp-servers
- https://strapi.io/blog/self-hosting-vs-managed-hosting
- https://datamcp.app/blog/hosted-vs-local-mcp-servers
- https://www.mcpbundles.com/pricing
- https://obot.ai/blog/hosted-mcp-vs-self-hosted-mcp-vs-local-mcp-pros-and-cons
- https://nimblebrain.ai/mcp/open-source-advantage/self-hosting-vs-managed







