I’d like to suggest an optional “Process when compute is available” mode for ChatGPT.

When sending a message, users could choose between:

  • Process now, normal response priority.
  • Process when compute is available, I’m happy to wait; process the request whenever capacity allows.

This would be different from Scheduled Tasks. The user wouldn’t specify when the task should run, they’d simply give OpenAI permission to process it whenever doing so is most efficient.

I think this could help OpenAI make better use of available compute, particularly for non-urgent tasks such as research, document analysis, coding, summarization, and other long-running requests.

As an incentive, deferred requests could potentially count less against usage limits or receive some other usage benefit. The user is effectively trading response speed for flexibility in when OpenAI uses its compute resources.

In short:

I don’t need my answer immediately → OpenAI gets more flexibility in scheduling compute → I receive a small usage benefit in return.

This could be completely optional, available to both free and paid users, and especially useful for people who are happy to let non-urgent workloads run whenever capacity is available.