> It's it bit slow since I'm not paying extra for ultrafast mode
Contrary to its name, "ultrafast" isn't faster than the rest, and many times slower than "max", as it'll "fork out" to a bunch of sub-agents and wait for them, + does extra "red-teaming" and more.
I think "ultrafast" is not referring to the speed of the "model" (harness in reality, as it's all the same model as "max") but rather how fast it consumes your usage limits.
> There is definitely a separate "Fast" mode that the app claims to yield a x1.5 speed increase (I have not tested it) with more token consumption.
Ah yes, I guess the portmanteau confused me and I assumed they were talking about "Ultra" the "reasoning effort" (which it isn't), rather than the "fast mode" which supposedly gives you priority over "non-fast mode requests". Although in practice, counter-intuitively, sometimes being in non-fast mode gives you faster replies than fast-mode, haven't got a feeling for why/when though.
He said "ultrafast" which has nothing to do with subagents. It's a new API tier where it runs on a different inference backend to get you faster token/s.
Contrary to its name, "ultrafast" isn't faster than the rest, and many times slower than "max", as it'll "fork out" to a bunch of sub-agents and wait for them, + does extra "red-teaming" and more.
I think "ultrafast" is not referring to the speed of the "model" (harness in reality, as it's all the same model as "max") but rather how fast it consumes your usage limits.