29. 09. 2026 Andrea Mariani AI, NetEye, Unified Monitoring

MCP or CLI? Pros, Cons, and the Real Limits with Local Models

More tools don’t mean more intelligence. Especially when the model runs on your own hardware.

Let’s be honest: ever since the Model Context Protocol became the de facto standard for connecting an LLM to the outside world, a sort of hoarding race has started. One MCP server for the filesystem, one for Git, one for Jira, one for the database, one for Icinga, one for the browser, one for email… and before long you find yourself with ten or fifteen servers attached to the same model. On paper it’s wonderful: the model “can do everything”. In practice, especially with a small-to-medium local model, it’s often the fastest way to end up with an assistant that’s slow, confused and unpredictable.

On the other side there’s a much older and far less fashionable approach: letting the model use the CLI. The same commands a sysadmin types every day, git, kubectl, curl, jq, icingacli, invoked by the model through a controlled shell.

In this post I’ll compare the two approaches without cheering for either side: what you gain and what you lose with MCP, what you gain and what you risk with the CLI and, above all, which concrete limits you run into when you connect many MCP servers to a 7-14 billion parameter model running locally. It’s the natural follow-up to what I wrote in Reflections on Running LLMs Locally and How to Build an LLM-Assisted Automation.

WHO THIS IS FORIf you already have a model running with Ollama, vLLM or similar, you’ve tried connecting a few MCP servers to it, and you’ve noticed that the more tools you add the worse it answers, this article is for you. You don’t need to be an ML engineer: you just need to have seen a tool call go wrong at least once.

1. What Are We Actually Talking About?

1.1 MCP in two lines

MCP is a standard protocol (JSON-RPC 2.0, stdio transport locally or HTTP remotely) through which a host application exposes a catalog of typed tools to the model. Each tool has a name, a natural-language description and a JSON schema for its parameters. The model reads the catalog, decides which tool to call, produces a JSON object with the parameters, and the host executes the call on the corresponding MCP server.

The key point, which we’ll come back to later: to be able to choose, the model must have the entire catalog in front of it. Every definition of every tool of every connected server ends up in the context, on every request.

1.2 The CLI as a “tool”

The CLI approach is blunter: the model gets a single tool, “run this command”, and writes the command line the way a person would. There’s no schema for each operation: there’s the knowledge the model already picked up during training from thousands of pages of documentation, man pages and Stack Overflow answers, plus whatever --help it reads on the spot.

These are two opposite philosophies: MCP describes to the model what it can do, the CLI trusts what the model already knows.

Be careful not to confuse this CLI with the agentic CLIs offered by the providers, such as Claude Code, Gemini CLI or Codex CLI. Those are the environment the agent runs in, not an alternative to MCP. As we’ll see in section 8, though, they’re exactly the ones pushing hardest for the “shell instead of tools” approach.

2. The Comparison in One Table

FeatureMCP serversCLI
Context cost− Every tool takes tokens, always+ Near zero: one generic tool
Input and output+ Typed, structured (JSON)− Free text, to be interpreted
Knowledge required from the model+ Low: the catalog explains everything− High: must know commands and flags
Hallucination risk+ Limited to names and parameters− Invented flags, wrong syntax
Composability− One call at a time+ Pipes, grep, jq, head
Access control+ Least privilege per server− A shell is potentially everything
Credential handling+ Encapsulated in the server− Environment variables, config files
Reproducible by a human− Requires an MCP client+ Copy the command and run it
Maintenance− A process and dependencies per server+ The binaries you already have
Suited to small models− Only with few tools− Only with well-known commands and a tight perimeter

As you can see, neither wins hands-down. Let’s look at the details.

3. MCP: Pros and Cons

3.1 The pros

  • A clear contract: every tool declares exactly what it accepts and what it returns. The model doesn’t have to guess the syntax, and the host can validate parameters before anything is even executed.
  • Structured output: results come back as JSON, ready to be verified by deterministic code without fragile text parsing.
  • Least privilege by design: the server that reads the CRM doesn’t write to the filesystem. Each server exposes only what’s needed, with its own credentials, isolated from the model.
  • No exposed shell: the model can’t “invent” a destructive command, because it can only call what exists in the catalog.
  • Standard and portable: the same server works with Open WebUI, with an IDE, with a custom orchestrator or with an agent in n8n.
  • Centralized governance: remote HTTP servers can sit behind a gateway with authentication, audit logging and rate limiting.

3.2 The cons

  • The catalog has a cost: every tool definition consumes tokens, and it does so on every single request, even when the tool isn’t needed.
  • Highly variable quality: many MCP servers are one-to-one wrappers of a REST API. The result: dozens of tools with similar names and generic descriptions that confuse the model.
  • Operational sprawl: every local server is a process, often launched with npx or uvx, with its own dependencies, versions and updates. Ten servers are ten things to maintain and monitor.
  • A new attack surface: a third-party server installed “because it was handy” is code running with your credentials. And a tool’s output (a web page, an email body, a ticket) can contain malicious instructions the model reads as if they were legitimate.
  • One call at a time: where a single pipe would do in a shell, MCP often needs several rounds of tool call → result → new tool call, each of which makes the context longer.

4. CLI: Pros and Cons

4.1 The pros

  • Knowledge already “included”: a decent model already knows how to use git log, systemctl status or curl. You don’t need to explain it with a JSON schema, and so you don’t need to spend tokens doing so.
  • Progressive discovery: if the model doesn’t know a command, it can read the --help only when needed. The context cost is paid on demand, not up front.
  • Composability: icingacli monitoring list services --problems | head -20 does in one line what would take several calls and a much larger output with MCP.
  • Filtering at the source: with grep, jq or awk the model can shrink the output before it enters the context. For a model with a small window, that’s a huge advantage.
  • Transparency: every action is a readable command. Whoever checks the logs sees exactly what was done and can reproduce it by hand.

4.2 The cons

  • A shell is a loaded gun: giving a model arbitrary command execution means giving it, potentially, everything the user it runs as can do. Without a sandbox and an allowlist it’s not an option, full stop.
  • Invented flags: small models in particular tend to “remember” options that don’t exist, or that exist in a different version of the tool. The command fails or, worse, does something different.
  • Unstructured output: space-aligned tables, localized messages, ANSI colors. Interpreting them is fragile, both for the model and for the verification code.
  • Interactive commands: a confirmation prompt, a pager, a password request: the process hangs and the agent gets stuck.
  • Environment dependency: different versions, different distributions, Windows versus Linux. The same command can behave differently on two machines.

5. The Real Problem: Many MCP Servers on a Small-to-Medium Local Model

And here we come to the point I care about most. With a frontier cloud model, with huge context windows and an excellent ability to choose among hundreds of tools, many of the problems below are mitigated. With a quantized 7-14B local model on one or two consumer GPUs, they become the factor that decides whether the system works or not.

5.1 The hidden cost: tool definitions

Every exposed tool is serialized into the prompt: name, description, parameter schema with types and descriptions. A well-written definition typically takes anywhere from about a hundred to a few hundred tokens; some popular servers expose dozens of tools each. Do the math and you’ll find that, before the user has typed a single word, a substantial part of the context window is already taken.

QUICK FORMULAOverhead tokens ≈ (number of servers) × (tools per server) × (average tokens per tool). Example: 8 servers × 12 tools × ~250 tokens ≈ 24,000 tokens. On a model configured with 8K or 16K of context, it simply doesn’t fit. On one with 32K, you’re left with very few tokens for the conversation, the tool results and the answer.

5.2 Context window and silent truncation

Here’s a trap I’ve seen many people fall into. Many models advertise large context windows (32K, 128K), but the serving engine doesn’t necessarily use them: in Ollama, for example, the context length actually allocated is governed by the num_ctx parameter, and its default has historically been much lower than what the model supports. When the prompt exceeds that limit, the input gets truncated. No obvious error: just a line in the logs.

The result is insidious: the system prompt or part of the tool definitions disappears, and the model starts behaving “strangely”, ignoring instructions or calling tools it can no longer see properly. Before blaming the model, always check the configured context:

FROM qwen2.5:14b
PARAMETER num_ctx 32768

or, per request via the API:

curl http://localhost:11434/api/chat -d '{
  "model": "qwen2.5:14b",
  "messages": [ ... ],
  "options": { "num_ctx": 32768 }
}'

5.3 More context means more VRAM

Raising num_ctx isn’t free. Every token in the context takes space in the KV cache, which lives in VRAM next to the model weights. For a model like Llama 3.1 8B with an FP16 KV cache, we’re talking about roughly 128 KB per token: 24,000 tokens of tool definitions alone are worth about 3 GB of VRAM. On a 16 or 24 GB GPU already busy holding the model weights, those gigabytes make the difference between a model fully on the GPU and one that partially spills into system RAM, with performance collapsing as a result.

5.4 Prefill and latency

Before generating the first token of its answer, the model has to process the entire prompt (the prefill phase). On a consumer GPU, a few tens of thousands of prompt tokens translate into seconds of waiting on every turn. And in an agent loop every tool call is a turn: five calls in a row mean five times that wait. Prefix caching in serving engines helps a lot, but only if the beginning of the prompt stays identical between requests: a tool catalog that changes order or content invalidates it.

5.5 Too much choice: tool selection degrades

Even when everything fits into the context, there’s a cognitive limit. A small model facing 100 tools chooses worse than it does facing 10. The typical mistakes:

  • Near-identical tools: get_services versus get_services_with_problems, search_issues versus list_issues. The model picks the “close” one instead of the right one.
  • Overlapping tools across servers: two servers that read files, two that do web searches. The model switches between them with no real criterion.
  • Forgotten tools: definitions buried in the middle of a long prompt get less attention (the so-called lost in the middle effect), and the right tool simply isn’t considered.
  • Pointless calls: the model calls a tool “just because”, since it sees it’s available, even when the answer doesn’t require it.

5.6 Malformed JSON and chat templates

Tool calling requires the model to produce syntactically valid JSON that conforms to the schema. Small models, especially with aggressive quantization, get it wrong more often: an extra comma, a missing required field, a wrong type, a slightly mangled tool name. On top of that there’s a lesser-known issue: tool calling depends on the model’s chat template, and a misaligned template (it happens with some GGUF conversions) can make behavior inconsistent even with a model that, on paper, supports tools natively.

5.7 Outputs that flood the context

The cost isn’t only on the way in. Many MCP servers return the full object from the underlying API: hundreds of fields, of which the model needs three. A list of 200 Icinga services in full JSON can be worth tens of thousands of tokens, and after two or three calls like that the context is saturated. Here the CLI with a | jq or a | head has a clear advantage.

5.8 Security: more servers, more surface

Every server you add is more code running with real credentials, and every tool that reads external content is a potential prompt injection channel. With a small model the risk grows: on average it’s worse at telling the user’s instructions apart from text contained in a tool result. If the same model has access both to a tool that reads email and to one that writes to a production system, you’ve built a bridge that someone, sooner or later, will try to cross.

SymptomLikely causeWhat to do
The model ignores the system promptTruncation due to insufficient num_ctxCheck the serving logs, raise the context or reduce the tools
Slow answers on every turnPrefill of a huge promptFewer tools, a stable catalog to benefit from prefix caching
It picks the wrong toolToo many similar or overlapping toolsPer-task profiles, distinguishable names and descriptions
Invalid JSON or wrong parametersSmall model, aggressive quantization, templateModels trained for tool use, higher quantization, validation and retry
Context saturated after a few callsOverly verbose tool outputsFilter and reduce output server-side
The model gets stuck in a loopUnhandled error, no limitPer-session call limit, clear errors returned to the model
Model partially in system RAMKV cache too large for the VRAMLess context, quantized KV cache, a larger GPU

6. How to Get Out of It: Practical Strategies

The good news is that none of these limits is a death sentence. Once again, it’s about moving the intelligence from the model to the scaffolding you build around it.

  • Few servers, chosen per task: don’t connect everything to everything. Define profiles: the “monitoring” profile has Icinga and the logs, the “development” profile has Git and the filesystem. The model only sees what the current task needs.
  • A tool router in front of the model: a preliminary step (an embedding search over tool descriptions, or a lightweight classifier) selects the 5-10 tools relevant to the request and passes only those to the model.
  • Coarse-grained tools: one well-built icinga_query tool beats twenty one-to-one API wrappers. Fewer choices, less ambiguity.
  • Short, distinguishable descriptions: each description should say in one sentence when to use that tool and when not to. Long details belong in documentation, not in the prompt.
  • Output filtered at the source: the server returns only the useful fields, with pagination and a size limit. The model should never receive a full dump.
  • The CLI behind a single MCP tool: the compromise I prefer. One tool, a strict allowlist of read-only commands and subcommands, a dedicated unprivileged system user, timeouts and output truncation. The model leverages what it already knows about the CLI, and you keep the kind of control MCP gives you.
  • A deterministic orchestrator: as described in How to Build an LLM-Assisted Automation, the model chooses, the code executes and verifies. The fewer decisions you ask of the model, the fewer tools it needs.

Here’s a minimal example of that compromise, written with FastMCP:

import shlex
import subprocess
from fastmcp import FastMCP

mcp = FastMCP("ops-cli")

# command -> allowed subcommands (None = no check on the subcommand)
ALLOWED = {
    "df": None,
    "uptime": None,
    "systemctl": {"status", "is-active"},
    "icingacli": {"monitoring"},
}
MAX_CHARS = 4000

@mcp.tool()
def run(command: str) -> str:
    """Run a read-only allowlisted command: df, uptime, systemctl status|is-active, icingacli monitoring."""
    argv = shlex.split(command)
    if not argv or argv[0] not in ALLOWED:
        return f"Command not allowed. Allowed: {', '.join(sorted(ALLOWED))}"
    subs = ALLOWED[argv[0]]
    if subs is not None and (len(argv) < 2 or argv[1] not in subs):
        return f"Subcommand not allowed for {argv[0]}. Allowed: {', '.join(sorted(subs))}"
    try:
        res = subprocess.run(argv, capture_output=True, text=True, timeout=30)
    except subprocess.TimeoutExpired:
        return "Command timed out after 30s"
    out = (res.stdout or res.stderr)[:MAX_CHARS]
    return f"exit={res.returncode}\n{out}"

if __name__ == "__main__":
    mcp.run()

Note three details: shell=False (the default for subprocess.run with an argument list), so no pipes and no chaining with ; or &&; a two-level allowlist, because systemctl on its own would also include stop and restart; and a limit on output size, because the context of a small model is the most precious resource you have. The process must run as a dedicated unprivileged user: the allowlist is one line of defense, not the only one.

QUICK TIPBefore adding a new MCP server, ask yourself: will the model use it in at least one request out of ten? If the answer is no, that server is paying rent in your context without doing any work. Keep it in a separate profile.

7. When to Choose What

ScenarioRecommended choiceWhy
Systems with rich APIs and sensitive credentials (CRM, ITSM, cloud)MCPIsolated credentials, least privilege, structured output
Read-only sysadmin operations on Linux hostsCLI with an allowlistThe model already knows the commands, filterable output
7-14B local model with a limited windowA few coarse-grained MCP tools + a controlled CLIMinimal context overhead
Non-technical users via chatMCP with profilesNo exposed shell, clear perimeter
Production automationsDeterministic orchestrator + few toolsPredictability, verification, audit
Personal exploration and prototypingSandboxed CLISpeed, zero configuration
HOW TO CHOOSESimple rule: use MCP where you need a contract (sensitive data, actions with side effects, non-technical users), use the CLI where you need efficiency (reads, diagnostics, well-known commands) and, with a small local model, keep the number of visible tools under ten in any case.

8. What About the Providers’ Agentic CLIs?

When people talk about “CLI” in the LLM world today, many think of something different from what we’ve described so far: the agentic CLIs the providers themselves promote, such as Anthropic’s Claude Code, Google’s Gemini CLI and OpenAI’s Codex CLI. It’s worth clarifying how they relate to the rest of this article.

8.1 They’re not an alternative to MCP: they’re a host

An agentic CLI is an application, exactly like Open WebUI or an IDE: it manages the conversation, the agent loop, permissions and tool execution. And it supports MCP, so you can connect the same servers we’ve been discussing. The right comparison isn’t “Claude Code versus MCP”, but “which tooling strategy does the agent adopt”.

8.2 The strategy they promote is precisely the “CLI” one

This is where the two meet. These CLIs give the agent a shell and encourage it to use command-line tools directly (git, gh, kubectl, aws) instead of one MCP server per service. On top of that come two mechanisms pointing in the same direction:

  • Skills and progressive disclosure: instead of loading every tool definition up front, the agent sees only a short index and loads a skill’s instructions and scripts only when the request calls for it. It’s the exact opposite of an MCP catalog that’s always in context.
  • Code instead of tool calls: the agent writes and runs a small script that calls the tools, filters the results and hands the model only what it needs. Anthropic itself has described this approach as a way to drastically cut token consumption compared with direct tool calls.

In other words, the providers have reached the same conclusion as this article: context is a precious resource, and a huge tool catalog is a poor way to spend it.

8.3 And with a local model?

Caution is needed here. Agentic CLIs are designed and optimized for their providers’ frontier models. Some can be pointed at a local model through a compatible endpoint, but with a small-to-medium model some specific limits emerge:

  • The starting cost is already high: the agent’s system prompt and built-in tool definitions take up a significant share of the context before you’ve added a single MCP server. With a 16K or 32K window, the margin shrinks quickly.
  • Long loops, repeated prefill: an agent that reads files, runs commands and checks results takes dozens of turns for a single task. On a consumer GPU, every turn pays the prefill of an ever-growing context.
  • High autonomy, lower reliability: these CLIs assume a model capable of planning many steps, correcting itself and knowing when to stop. A 7-14B model tends to get lost, repeat actions or declare a job finished when it isn’t.
  • A real shell: everything in section 4.2 applies twice over. The CLI’s permission system helps, but a less reliable model makes sandboxing and confirmation of risky actions even more important.
AspectAgentic CLI (shell + skills)Generic host + many MCP servers
Initial context overhead− Substantial agent system prompt− Grows with every connected server
Overhead growth+ Contained: skills load on demand− Linear with the number of tools
Autonomy required from the model− High: plans and executes many steps+ Can be constrained by an orchestrator
Optimized forFrontier cloud modelsAny model, if the catalog is small
With a 7-14B local model− Usable for short, well-bounded tasks− Usable only with few tools
IN PRACTICEIt’s worth borrowing the ideas behind agentic CLIs, even if you don’t use them: a few generic primitives instead of dozens of tools, instructions loaded only when needed, output filtered by code before it reaches the model. These are exactly the strategies from section 6, and they work with a local model inside a deterministic orchestrator too.

Conclusions

MCP and the CLI aren’t rivals: they’re two tools with different trade-offs. MCP gives you structure, security and a clear contract, but you pay for it in tokens, VRAM and latency, and the bill grows with every server you connect. The CLI is efficient and leverages what the model already knows, but without a strict perimeter it’s a risk no production environment should accept.

With a small-to-medium local model, the lesson is always the same: the model doesn’t become more capable because you give it more tools. It becomes more capable when you give it the right tools, few and well described, and when the heavy lifting is done by the code around it. Everything else is wasted context.

Further Reading

Andrea Mariani

Andrea Mariani

Author

Andrea Mariani

Leave a Reply

Your email address will not be published. Required fields are marked *

Archive