At first I was mislead by some faulty benchmarking I was doing and thought that llama.cpp + SYCL was worse than Vulkan stock. Vulkan stock couldn't run past 10t/s with MTP for some reason, so after a whole bunch of shenanigans I ended up again back on llama.cpp + SYCL but was lead to believe MTP was hurting me so I didn't retry until last night. I finally got the 35-40t/s with llama.cpp + SYCL and MTP and 800PP with AOT over using JIT. I drop to 25ish after 40-50k context and stay kinda flat 12ish at 160k
I didn't want to move to Linux to try vLLM so this was all just trying to fight windows b.s.
If you're on windows and fighting with the VRAM offload after 70 seconds like I was: HKLM\SYSTEM\CurrentControlSet\Control\GraphicDrivers
EnableRuntimePowerManagement set to 0