Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Not directly related to K2, but why do a lot of the newly released models basically say day zero day support in vllm, slang but often not llama.cpp?

Llama.cpp is then often a few days behind, which given it's the only inference engine supporting older architectures is quite frustrating.

 help



Developers with lots of VC money to burn are working on things like B100/B200/B300 which are well supported in VLLM, everything else in terms of supporting more mundane GPUs or other platforms is ancillary to the main task of getting the thing trained and aligned.



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: