Developers with lots of VC money to burn are working on things like B100/B200/B300 which are well supported in VLLM, everything else in terms of supporting more mundane GPUs or other platforms is ancillary to the main task of getting the thing trained and aligned.
Llama.cpp is then often a few days behind, which given it's the only inference engine supporting older architectures is quite frustrating.