Moving Beyond General-Purpose Vision Language Models
Merve Noyan argues that developers should transition from directly using Vision Language Models (VLMs) to building end-to-end vision applications that leverage specialized skills.
She notes that excessive reliance on large VLMs often fails to deliver real-time performance, frequently yielding only 30-40 frames per second.
Furthermore, dedicated models like RT-DETR consistently demonstrate superior performance in specific tasks, while restrictive licenses such as AGPL 3.0, used in models like YOLO, necessitate a shift toward Apache 2.0 alternatives for broader adoption.
"I don't want developers to directly use vision language models anymore. And uh I want every single developer to start uh building vision langu vision applications end to end."
Merve Noyan


