Hi all,
I regularly set Codex on well scoped work in the evening but last night things went haywire. My Codex run consumed the remainder of my weekly pro plan allowance (80%+) and +$700 of additional usage credits before my spending cap finally stopped it. Effectively zero work was done.
Codex’s own subsequent postmortem acknowledged that:
- 35 queue packets in one night, 13 of them READY drops. Thirteen preregistration package directories for what is nominally one lease — nine numbered revisions plus draft/staging copies, 568 KB of seal material, most of it re-derived content.
- Recorded scope moved from 35 to 66 to 69 to 73 to 112 paths.
- The root cause: scope discovery and implementation were interleaved. Codex used a scanner it was simultaneously fixing to define the boundary of its own lease. Every scanner improvement invalidated the previous boundary, and each invalidation cost a full re-census of every path, a fresh seal package, a queue round trip, and a complete re-verification from me — 74 blob hashes across the amendments, plus a full seal read each time. Five scope moves, nine revisions.
I understand that I authorized the original task. I am not claiming that my account or payment method was compromised. However, I did not authorize or reasonably expect an uncontrolled recursive workflow that consumed nearly an entire weekly allowance plus ~$700 while I was asleep.
Has anyone dealt with this before and thoughts on additional controls you put in place to mitigate this risk?
Welcome to the forum!
I can not say this will help with your specific problem but it has worked for me and the other user who also learned from it.
I think many people in this community will recognize this scenario. A large AGENTS.md. Maximum context. A detailed prompt. One big /goal. Codex starts working confidently: it explores the repository, builds a plan, and changes the code. Then the scope expands, unnecessary files are read, one fix creates another problem, and in the end we get an expensive result that still needs to be partially rewritten. For a long time, I thought the solution was to give the model even more information an…
A useful lesson from today. I had a business idea that required collecting, filtering, enriching, and analyzing a large number of small-business records. I initially prepared a very detailed prompt and ran the entire task through an OpenAI reasoning model at maximum settings. The result: all available tokens were consumed in about 14 minutes, while the task itself was still far from complete. Then I followed advice that another developer from this community had given me earlier: repetitive an…
So, unless you tell Sol to stop it will do that. If you aren’t using Sol, think again. Unless you’ve chosen 5.6 or kept reasoning to low, you’re now using Sol.