Show HN: Cuts Long Horizon Inference Costs by 50% via external KV Cache Offload

(github.com)

21 points | by arnav__1 8 hours ago ago

1 comments