Agentic coding tools promise to move beyond code completion. Instead of asking an assistant to generate a function or explain an error, we give it a goal and let it inspect a repository, plan a change, edit several files, run tests, and iterate.
That shift is powerful, but it also changes the engineering trade-offs: success depends not only on which model produces the best code, but also on how much context and orchestration a task needs, which models should handle different roles, how often verification should run, and how much of the token budget is spent on useful work rather than coordination.
I started exploring these questions with OpenCode.
My first choice was OpenCode because it is open source and does not force the workflow into a single model provider. Its provider-agnostic approach makes it possible to experiment with different models for different kinds of work: a strong reasoning model for architecture, a fast and inexpensive model for repository exploration, and another model for implementation or review.
That flexibility matters in real development. A coding workflow is not one homogeneous task. Repository search, implementation, debugging, documentation lookup, UI work, and final review have different requirements. Using the same expensive model for all of them is convenient, but it is not necessarily efficient.
OpenCode also provides the practical foundation expected from a modern coding agent, including terminal usage, editor integrations, LSP support, and multiple sessions. The important point for me, however, was the separation between the agent interface and the model provider. I could change the model strategy without abandoning the tool that manages the coding session.
This makes OpenCode a useful base layer for experimentation. It is not only a coding assistant; it is also a place to test different approaches to agent orchestration.
The next step was Oh My OpenAgent, also presented by the project as OmO. It turns OpenCode into a much more opinionated multi-agent system.
The basic idea is attractive: instead of relying on one general-purpose agent, give different responsibilities to specialized agents and let an orchestrator decide how to combine them. The project describes roles such as planning, exploration, research, implementation, design, and verification. It also adds mechanisms for parallel work, background tasks, model mixing, skills, memory, and tool-call coordination.
OmO gives complex tasks an explicit division of labor. Specialized agents can explore the repository, advise on architecture, implement changes, research unfamiliar areas, and review the result. I found this particularly useful for large features, migrations, and architectural refactors.
Its model routing also helps when configured carefully: fast or inexpensive models can handle reconnaissance and mechanical work, while stronger models focus on architecture, debugging, or review. Parallel investigations keep the primary session focused, and the workflow encourages a complete loop of implementation, testing, and verification. Memory and reusable skills can also carry project conventions across sessions.
These benefits are task-dependent. OmO was most useful when the work was broad enough to justify specialization and coordination.
My main concern with OmO was token consumption. It often felt bloated and used considerably more context than vanilla OpenCode. Delegated tasks require their own instructions and context, while coordination, repeated repository scans, planning, testing, and review add further model calls.
I also found it difficult to keep OmO in a direct coding mode. It could delegate work that would have been faster to complete directly, add latency, or continue investigating after a practical answer was available. The many agents, prompts, model assignments, and configuration choices made the setup harder to understand and tune.
This overhead may be worthwhile for migrations and large refactors, but it was difficult to justify for small bug fixes. For me, OmO is a powerful workflow for selected tasks rather than a universal replacement for vanilla OpenCode.
I then tried Oh My OpenCode Slim. Although it began as a fork, OmO-Slim has evolved into a distinct project with its own orchestration model, agent roles, and trade-offs. Its design centers on a focused group of specialists working through an orchestrator.
The current project describes six main roles:
OmO-Slim felt more deliberate and easier to control in my workflow. Its defined roles made delegation easier to understand and reduced the amount of orchestration around each task.
The configurable model assignments were particularly useful. I could use inexpensive, fast models for exploration or implementation and reserve stronger models for architecture, debugging, or review. In practice, this gave me a better balance of capability, speed, and token usage while preserving the benefits of delegated work.
Background tasks also made the workflow smoother. I could keep interacting with the orchestrator while other agents worked.
The improvement was relative, not absolute: OmO-Slim used fewer tokens than OmO in my experience, but it was still far from lightweight. Auto-delegation, parallel agents, separate instructions, and repeated context all added to the bill, especially for tasks that did not really need a team of agents.
The setup also requires tuning. Model assignments and enabled capabilities affect both the result and the token usage, while a misunderstanding by the orchestrator can send several delegated agents in the wrong direction. For me, OmO-Slim is a more practical compromise than OmO, but it remains a token-consuming workflow rather than a free efficiency gain.
The main lesson from both experiments is that orchestration has a cost. It can be valuable for large refactors, but for smaller tasks the extra delegation, context, and coordination may not justify the result. Parallel agents can also repeat the same investigation or create overlapping work. I am therefore still looking for a workflow that adapts its level of orchestration to the task, instead of applying the same process every time.
There are good reasons to evaluate proprietary tools as well as open ones. A more integrated product may offer a smoother workflow, stronger defaults, or a different balance between autonomy, supervision, and token usage. I want to compare those trade-offs rather than rule out a category in advance.
I am already trying Claude Code by Anthropic, and my first impressions are positive. It represents a more integrated and opinionated approach to agentic coding, with the product experience, prompts, and agent behavior designed as one system.
I am still evaluating whether it can complete common development tasks with less token overhead and reduce the supervision I need to provide. The trade-off is a more limited provider strategy, but the workflow has been promising enough to deserve a serious comparison.
I have also tried Codex from OpenAI, although only briefly so far. It offers another point of comparison for autonomous coding workflows and may help separate the benefits of agentic execution from the benefits of a particular OpenCode orchestration design.
I want to evaluate it more carefully on repository exploration, multi-file changes, test execution, and recovery from failed attempts before drawing any conclusions.
My preferred non-proprietary path is to stay with OpenCode and configure it directly rather than add another orchestration framework. Custom agents, focused prompts, and a small number of explicit tools may provide enough structure for many projects while preserving control over context, delegation, and model choice.
This is the setup I want to explore thoroughly before making a final choice. Any more complex approach should offer enough additional value to justify its extra prompts, model calls, and operational complexity.
I also want to explore other emerging harnesses out of personal curiosity. I want to understand how different projects approach autonomy, extensibility, model choice, and orchestration, even if I ultimately stay with OpenCode. This exploration may also reveal which parts of my workflow should belong in the harness itself and which are better implemented as custom agents, prompts, or small tools.
I do not need a large benchmark to form a useful first impression. For whichever tool I try, I want to use the same repository and a few representative tasks: small tasks, bug fixes, and larger multi-file features or refactors.
I will look at more than token usage. The output quality matters, as does the interaction with the model: how clearly it understands the task, how easy it is to guide, how often it needs correction, and whether the workflow feels natural. Time taken and the amount of supervision required are also part of the picture.
I am still looking for a harness that adds useful autonomy without getting in the way. The best option for me will not necessarily be the one with the most agents or the lowest token usage, but the one that produces good results and makes the development process feel better.
Project capabilities and links checked on 30 September 2026. Token-usage observations in this article are personal experience rather than a controlled benchmark.
Did you find this article interesting? Are you an “under the hood” kind of person? We’re really big on automation and we’re always looking for people in a similar vein to fill roles like this one as well as other roles here at Würth IT Italy.