
Kog is hiring a
GPU Engineer
About Kog
Kog builds high-performance AI inference software, delivering the fastest LLM inference engine on standard datacenter GPUs. We focus on low-level kernel work, monokernel pipelines, and autonomous optimization of kernels for large language models.
Job Description
We are hiring a GPU Engineer to work on the fastest LLM inference engine on standard datacenter GPUs. You will own low-level kernel work in CUDA/PTX or HIP/CDNA ISA, the monokernel pipeline, profiling infrastructure inside it, scaling to frontier MoE models that run in production, and building our own agents that optimize kernels and inference autonomously. We generate 3,000 tokens/s per request on 8x AMD MI300X and 2,100 on 8x NVIDIA H200 at batch size 1, FP16 with no speculative decoding. At batch size 1 the decode is GEMV, memory bandwidth bound, and MBU counts. We rewrote the hot path ourselves from assembly on the chip up to the Transformer we designed around it, with the full decode running as a single persistent GPU kernel. See playground.kog.ai. Showing your code is part of the process. If you are outside a Europe-compatible timezone, relocation to one is required. Remote within Europe-compatible timezones; onsite in Paris one week per month. Apply: https://jobs.ashbyhq.com/kog/e3950334-a2a6-43cc-a744-df6c38683166. Questions, email nicolas.constant@kog.aiLoading...
Share this job
Hiring engineers?
Reach thousands of tech candidates from the Hacker News community.
Post a Job β $99










