I am no expert and i still need to perform some experiments but i believe skills and plugins can increase “usage burn” .
For what i understand, usage is measured by Input, cached input and output tokens. Each agent spawned will have their own model+reasoning efford and token usage. Also, Sol-Ultra can spawn extra agents also.
At the end of the day, if something we are working on is running on a Sol-High or xHigh (or Max) and we allow/instruct for agents to run, if those agents also run at high reasoning cost, our usage will burn way faster.

On my side, i’ve also been using, until this week, a plugin called “Superpowers” to help with my vibing ^_^. this plugin spawns a crazy amount of agents . The result is good but the cost is to great.

What i decided to do was to stop using plugins that i believe have low ROI and start using more config and agents.md to give codex the necessary tools to spawn agents with a more efficient model-effort assignment. Example of part of what i wrote on agents.md for suggested escalation:

  • Luna low/medium: mechanical, deterministic, repetitive, search-heavy, extraction, simple test/log analysis, and clearly specified work
  • Terra medium: normal software engineering and implementation.
  • Terra high/xhigh: difficult debugging, non-trivial review, complex logic,
    integration work, and substantial test reasoning.
  • Sol medium/high: ambiguous cross-cutting implementation, difficult diagnosis,
    architecture, or decisions requiring substantial judgment.
  • Sol xhigh/max: reserve for exceptionally difficult, high-risk, or unresolved
    problems where additional reasoning is likely to materially improve the result.

This way, even if main model prompt is running on Sol-xHigh (because why not.. :smiley: ) it can spawn an explorer running on Luna Medium, then have couple workers, one on Luna Medium and another on Tera Medium because it needs more juice.

Another thing, agent spawn is reasoned to not happen if the task can be performed by the main session model. For what i can tell, if for each agent we need to give it enough context for it to act, and then it also needs to retrieve, for example, repo information, then we increase usage count unnecessarily.

I might be “off” in some of my assumptions (again, not an expert ^_^) but, at least in theory, they make sense. I just need a fresh reset to test this out because, well, i’m at 0% right now :smiley: