VR
VLM Run is hiring a
Founding Infrastructure Engineer
About VLM Run
VLM Run is a first-of-its-kind API dedicated to running Vision Language Models on Documents, Images, and Video. We’re building a stack from the bottom-up for ‘Visual’ applications of language models that we believe will make up > 90% of inference needs in the next 5 years.
Job Description
We’re building the inference platform for visual intelligence. This role owns the infrastructure problem for the VLM Run Gateway, serving open-weight VLMs, embodied VLAs, ViTs, across GPUs and clouds. The team is highly technical and has shipped production ML infra in autonomous driving and LLMs. This is a founding, in-person role in the Bay Area with the goal to scale to 100s of thousands of requests per day. Applicants should email hiring@vlm.run with their GitHub profile, papers, and ML projects (100+ GH stars) that demonstrate experience with Ray, Kubernetes, GPUs, and serverless. No remote work; must be in the Bay Area. Short emails get faster responses.Remote Conditions
No remote work; must be in the Bay Area.
Location
Bay Area, CA
Salary
Not Specified
Benefits
Not Specified
Tech Tags
Senior Role
Date Listed
01 October, 2026 (about 14 hours ago)










