by Beefin on 1/26/26, 7:35 PM with 1 comments
by storystarling on 1/26/26, 9:13 PM
What hardware are you running this on to get 2-3s latency? A 14GB model plus KV cache seems like it would require a 24GB card (3090/4090) to avoid swapping. I've found that once you spill over to system RAM on consumer gear the performance usually falls off a cliff.