More tools don’t mean more intelligence. Especially when the model runs on your own hardware.
Let’s be honest: ever since the Model Context Protocol became the de facto standard for connecting an LLM to the outside world, a sort of hoarding race has started. One MCP server for the filesystem, one for Git, one for Jira, one for the database, one for Icinga, one for the browser, one for email… and before long you find yourself with ten or fifteen servers attached to the same model. On paper it’s wonderful: the model “can do everything”. In practice, especially with a small-to-medium local model, it’s often the fastest way to end up with an assistant that’s slow, confused and unpredictable.
On the other side there’s a much older and far less fashionable approach: letting the model use the CLI. The same commands a sysadmin types every day, git, kubectl, curl, jq, icingacli, invoked by the model through a controlled shell.
In this post I’ll compare the two approaches without cheering for either side: what you gain and what you lose with MCP, what you gain and what you risk with the CLI and, above all, which concrete limits you run into when you connect many MCP servers to a 7-14 billion parameter model running locally. It’s the natural follow-up to what I wrote in Reflections on Running LLMs Locally and How to Build an LLM-Assisted Automation.
| WHO THIS IS FOR | If you already have a model running with Ollama, vLLM or similar, you’ve tried connecting a few MCP servers to it, and you’ve noticed that the more tools you add the worse it answers, this article is for you. You don’t need to be an ML engineer: you just need to have seen a tool call go wrong at least once. |
1.1 MCP in two lines
MCP is a standard protocol (JSON-RPC 2.0, stdio transport locally or HTTP remotely) through which a host application exposes a catalog of typed tools to the model. Each tool has a name, a natural-language description and a JSON schema for its parameters. The model reads the catalog, decides which tool to call, produces a JSON object with the parameters, and the host executes the call on the corresponding MCP server.
The key point, which we’ll come back to later: to be able to choose, the model must have the entire catalog in front of it. Every definition of every tool of every connected server ends up in the context, on every request.
1.2 The CLI as a “tool”
The CLI approach is blunter: the model gets a single tool, “run this command”, and writes the command line the way a person would. There’s no schema for each operation: there’s the knowledge the model already picked up during training from thousands of pages of documentation, man pages and Stack Overflow answers, plus whatever --help it reads on the spot.
These are two opposite philosophies: MCP describes to the model what it can do, the CLI trusts what the model already knows.
Be careful not to confuse this CLI with the agentic CLIs offered by the providers, such as Claude Code, Gemini CLI or Codex CLI. Those are the environment the agent runs in, not an alternative to MCP. As we’ll see in section 8, though, they’re exactly the ones pushing hardest for the “shell instead of tools” approach.
| Feature | MCP servers | CLI |
|---|---|---|
| Context cost | − Every tool takes tokens, always | + Near zero: one generic tool |
| Input and output | + Typed, structured (JSON) | − Free text, to be interpreted |
| Knowledge required from the model | + Low: the catalog explains everything | − High: must know commands and flags |
| Hallucination risk | + Limited to names and parameters | − Invented flags, wrong syntax |
| Composability | − One call at a time | + Pipes, grep, jq, head |
| Access control | + Least privilege per server | − A shell is potentially everything |
| Credential handling | + Encapsulated in the server | − Environment variables, config files |
| Reproducible by a human | − Requires an MCP client | + Copy the command and run it |
| Maintenance | − A process and dependencies per server | + The binaries you already have |
| Suited to small models | − Only with few tools | − Only with well-known commands and a tight perimeter |
As you can see, neither wins hands-down. Let’s look at the details.
3.1 The pros
3.2 The cons
npx or uvx, with its own dependencies, versions and updates. Ten servers are ten things to maintain and monitor.4.1 The pros
git log, systemctl status or curl. You don’t need to explain it with a JSON schema, and so you don’t need to spend tokens doing so.--help only when needed. The context cost is paid on demand, not up front.icingacli monitoring list services --problems | head -20 does in one line what would take several calls and a much larger output with MCP.grep, jq or awk the model can shrink the output before it enters the context. For a model with a small window, that’s a huge advantage.4.2 The cons
And here we come to the point I care about most. With a frontier cloud model, with huge context windows and an excellent ability to choose among hundreds of tools, many of the problems below are mitigated. With a quantized 7-14B local model on one or two consumer GPUs, they become the factor that decides whether the system works or not.
5.1 The hidden cost: tool definitions
Every exposed tool is serialized into the prompt: name, description, parameter schema with types and descriptions. A well-written definition typically takes anywhere from about a hundred to a few hundred tokens; some popular servers expose dozens of tools each. Do the math and you’ll find that, before the user has typed a single word, a substantial part of the context window is already taken.
| QUICK FORMULA | Overhead tokens ≈ (number of servers) × (tools per server) × (average tokens per tool). Example: 8 servers × 12 tools × ~250 tokens ≈ 24,000 tokens. On a model configured with 8K or 16K of context, it simply doesn’t fit. On one with 32K, you’re left with very few tokens for the conversation, the tool results and the answer. |
5.2 Context window and silent truncation
Here’s a trap I’ve seen many people fall into. Many models advertise large context windows (32K, 128K), but the serving engine doesn’t necessarily use them: in Ollama, for example, the context length actually allocated is governed by the num_ctx parameter, and its default has historically been much lower than what the model supports. When the prompt exceeds that limit, the input gets truncated. No obvious error: just a line in the logs.
The result is insidious: the system prompt or part of the tool definitions disappears, and the model starts behaving “strangely”, ignoring instructions or calling tools it can no longer see properly. Before blaming the model, always check the configured context:
FROM qwen2.5:14b
PARAMETER num_ctx 32768
or, per request via the API:
curl http://localhost:11434/api/chat -d '{
"model": "qwen2.5:14b",
"messages": [ ... ],
"options": { "num_ctx": 32768 }
}'
5.3 More context means more VRAM
Raising num_ctx isn’t free. Every token in the context takes space in the KV cache, which lives in VRAM next to the model weights. For a model like Llama 3.1 8B with an FP16 KV cache, we’re talking about roughly 128 KB per token: 24,000 tokens of tool definitions alone are worth about 3 GB of VRAM. On a 16 or 24 GB GPU already busy holding the model weights, those gigabytes make the difference between a model fully on the GPU and one that partially spills into system RAM, with performance collapsing as a result.
5.4 Prefill and latency
Before generating the first token of its answer, the model has to process the entire prompt (the prefill phase). On a consumer GPU, a few tens of thousands of prompt tokens translate into seconds of waiting on every turn. And in an agent loop every tool call is a turn: five calls in a row mean five times that wait. Prefix caching in serving engines helps a lot, but only if the beginning of the prompt stays identical between requests: a tool catalog that changes order or content invalidates it.
5.5 Too much choice: tool selection degrades
Even when everything fits into the context, there’s a cognitive limit. A small model facing 100 tools chooses worse than it does facing 10. The typical mistakes:
get_services versus get_services_with_problems, search_issues versus list_issues. The model picks the “close” one instead of the right one.5.6 Malformed JSON and chat templates
Tool calling requires the model to produce syntactically valid JSON that conforms to the schema. Small models, especially with aggressive quantization, get it wrong more often: an extra comma, a missing required field, a wrong type, a slightly mangled tool name. On top of that there’s a lesser-known issue: tool calling depends on the model’s chat template, and a misaligned template (it happens with some GGUF conversions) can make behavior inconsistent even with a model that, on paper, supports tools natively.
5.7 Outputs that flood the context
The cost isn’t only on the way in. Many MCP servers return the full object from the underlying API: hundreds of fields, of which the model needs three. A list of 200 Icinga services in full JSON can be worth tens of thousands of tokens, and after two or three calls like that the context is saturated. Here the CLI with a | jq or a | head has a clear advantage.
5.8 Security: more servers, more surface
Every server you add is more code running with real credentials, and every tool that reads external content is a potential prompt injection channel. With a small model the risk grows: on average it’s worse at telling the user’s instructions apart from text contained in a tool result. If the same model has access both to a tool that reads email and to one that writes to a production system, you’ve built a bridge that someone, sooner or later, will try to cross.
| Symptom | Likely cause | What to do |
|---|---|---|
| The model ignores the system prompt | Truncation due to insufficient num_ctx | Check the serving logs, raise the context or reduce the tools |
| Slow answers on every turn | Prefill of a huge prompt | Fewer tools, a stable catalog to benefit from prefix caching |
| It picks the wrong tool | Too many similar or overlapping tools | Per-task profiles, distinguishable names and descriptions |
| Invalid JSON or wrong parameters | Small model, aggressive quantization, template | Models trained for tool use, higher quantization, validation and retry |
| Context saturated after a few calls | Overly verbose tool outputs | Filter and reduce output server-side |
| The model gets stuck in a loop | Unhandled error, no limit | Per-session call limit, clear errors returned to the model |
| Model partially in system RAM | KV cache too large for the VRAM | Less context, quantized KV cache, a larger GPU |
The good news is that none of these limits is a death sentence. Once again, it’s about moving the intelligence from the model to the scaffolding you build around it.
icinga_query tool beats twenty one-to-one API wrappers. Fewer choices, less ambiguity.Here’s a minimal example of that compromise, written with FastMCP:
import shlex
import subprocess
from fastmcp import FastMCP
mcp = FastMCP("ops-cli")
# command -> allowed subcommands (None = no check on the subcommand)
ALLOWED = {
"df": None,
"uptime": None,
"systemctl": {"status", "is-active"},
"icingacli": {"monitoring"},
}
MAX_CHARS = 4000
@mcp.tool()
def run(command: str) -> str:
"""Run a read-only allowlisted command: df, uptime, systemctl status|is-active, icingacli monitoring."""
argv = shlex.split(command)
if not argv or argv[0] not in ALLOWED:
return f"Command not allowed. Allowed: {', '.join(sorted(ALLOWED))}"
subs = ALLOWED[argv[0]]
if subs is not None and (len(argv) < 2 or argv[1] not in subs):
return f"Subcommand not allowed for {argv[0]}. Allowed: {', '.join(sorted(subs))}"
try:
res = subprocess.run(argv, capture_output=True, text=True, timeout=30)
except subprocess.TimeoutExpired:
return "Command timed out after 30s"
out = (res.stdout or res.stderr)[:MAX_CHARS]
return f"exit={res.returncode}\n{out}"
if __name__ == "__main__":
mcp.run()
Note three details: shell=False (the default for subprocess.run with an argument list), so no pipes and no chaining with ; or &&; a two-level allowlist, because systemctl on its own would also include stop and restart; and a limit on output size, because the context of a small model is the most precious resource you have. The process must run as a dedicated unprivileged user: the allowlist is one line of defense, not the only one.
| QUICK TIP | Before adding a new MCP server, ask yourself: will the model use it in at least one request out of ten? If the answer is no, that server is paying rent in your context without doing any work. Keep it in a separate profile. |
| Scenario | Recommended choice | Why |
|---|---|---|
| Systems with rich APIs and sensitive credentials (CRM, ITSM, cloud) | MCP | Isolated credentials, least privilege, structured output |
| Read-only sysadmin operations on Linux hosts | CLI with an allowlist | The model already knows the commands, filterable output |
| 7-14B local model with a limited window | A few coarse-grained MCP tools + a controlled CLI | Minimal context overhead |
| Non-technical users via chat | MCP with profiles | No exposed shell, clear perimeter |
| Production automations | Deterministic orchestrator + few tools | Predictability, verification, audit |
| Personal exploration and prototyping | Sandboxed CLI | Speed, zero configuration |
| HOW TO CHOOSE | Simple rule: use MCP where you need a contract (sensitive data, actions with side effects, non-technical users), use the CLI where you need efficiency (reads, diagnostics, well-known commands) and, with a small local model, keep the number of visible tools under ten in any case. |
When people talk about “CLI” in the LLM world today, many think of something different from what we’ve described so far: the agentic CLIs the providers themselves promote, such as Anthropic’s Claude Code, Google’s Gemini CLI and OpenAI’s Codex CLI. It’s worth clarifying how they relate to the rest of this article.
8.1 They’re not an alternative to MCP: they’re a host
An agentic CLI is an application, exactly like Open WebUI or an IDE: it manages the conversation, the agent loop, permissions and tool execution. And it supports MCP, so you can connect the same servers we’ve been discussing. The right comparison isn’t “Claude Code versus MCP”, but “which tooling strategy does the agent adopt”.
8.2 The strategy they promote is precisely the “CLI” one
This is where the two meet. These CLIs give the agent a shell and encourage it to use command-line tools directly (git, gh, kubectl, aws) instead of one MCP server per service. On top of that come two mechanisms pointing in the same direction:
In other words, the providers have reached the same conclusion as this article: context is a precious resource, and a huge tool catalog is a poor way to spend it.
8.3 And with a local model?
Caution is needed here. Agentic CLIs are designed and optimized for their providers’ frontier models. Some can be pointed at a local model through a compatible endpoint, but with a small-to-medium model some specific limits emerge:
| Aspect | Agentic CLI (shell + skills) | Generic host + many MCP servers |
|---|---|---|
| Initial context overhead | − Substantial agent system prompt | − Grows with every connected server |
| Overhead growth | + Contained: skills load on demand | − Linear with the number of tools |
| Autonomy required from the model | − High: plans and executes many steps | + Can be constrained by an orchestrator |
| Optimized for | Frontier cloud models | Any model, if the catalog is small |
| With a 7-14B local model | − Usable for short, well-bounded tasks | − Usable only with few tools |
| IN PRACTICE | It’s worth borrowing the ideas behind agentic CLIs, even if you don’t use them: a few generic primitives instead of dozens of tools, instructions loaded only when needed, output filtered by code before it reaches the model. These are exactly the strategies from section 6, and they work with a local model inside a deterministic orchestrator too. |
MCP and the CLI aren’t rivals: they’re two tools with different trade-offs. MCP gives you structure, security and a clear contract, but you pay for it in tokens, VRAM and latency, and the bill grows with every server you connect. The CLI is efficient and leverages what the model already knows, but without a strict perimeter it’s a risk no production environment should accept.
With a small-to-medium local model, the lesson is always the same: the model doesn’t become more capable because you give it more tools. It becomes more capable when you give it the right tools, few and well described, and when the heavy lifting is done by the code around it. Everything else is wasted context.
num_ctx parameter