Yes, expression challenges had some potential upside.
Did you notice that caching appears a bit volatile? Sometimes nothing is cached at all.
I am running one model-effort combination at a time repeatedly, and sometimes cache hits grow, while other times there is literally nothing cached across different combinations.
coming soon, multi session monitoring:
That was pushed and includes additional telemetry.
Now I’ve just merged improvements to the benchmarks tab:
- ranking on the benchmarks tab, you can rank prioritising cost or speed (but taking into account the other rank) - review the README for more info
- simplified some options (there’s now only one collection of benchmarks)
You could already order by time and cost so this allows you to blend those two in different ways.
I’ve also improved the interface to gather together the different quota views in one tab where you can select the specific style of quota view you prefer and enhancements to the status at the top to reflect quota status so it’s clear no matter what tab you are on:
this quota status takes into account current known reset dates so shouldn’t flag if reset date and rate of consumption are reasonable.
All these resets are making my UI look boring!
In any case, lets get motoring people!
I’ve added experimental quota pricing estimation.
After a while of using quota whilst codexometer is running, it will attempt to estimate your “API Equivalent” spend and also what your potential max api equivalent spend might be.
Regard this as an educated estimate and not foolproof.
The explanation of how we arrive at the figures is here: GitHub - merefield/codexometer: A terminal widget that allows you to keep track of Codex usage against your current quota · GitHub with appropriate disclaimers.