As the title says - on a mathematical-related query without agentic calls at 128K max tokens context for GPT-5.6 Sol on the API Platform (not ChatGPT/Codex) for a single query:
```json
{
“model”: “gpt-5.6-sol”,
“max_output_tokens”: 128000,
“service_tier”: “flex”,
“background”: true,
“reasoning”: {“effort”: “max”, “summary”: “auto”},
“tools”: []
}
```
I was charged for 16M output tokens (reasoning tokens correctly stopped at 128K). What is even more concerning is that the model shouldn’t support that amount of output tokens for a single query at all. This, of course, cost $ 250 and was replicated through several hundred calls on different mathematical questions, racking up 1000s of dollars in costs. I contacted the support, and they are completely unhelpful - I only get LLM answers about checking the documentation of Codex despite explaining 3 times so far that this happened on the API Platform and providing exact information like key, org, response ids etc. I am super frustrated with the support, and I do not know what to do. Additionally, this has been dragging for days now, and we are left holding the bag. Please advise what to do here!
Similar issue, also unresolved: Gpt-5.6: usage.output_tokens is ~9x actual generation (re-summed once per reasoning item), exceeds max_output_tokens, and is what gets billed
And yes, the support responses were nonsensical, simply pointing generically at unrelated help articles.
Very frustrating.
They closed all of my support tickets … I even disputed one and they closed the dispute too. I also send them your ticket last week and nothing… We are a bit out of ideas what to do