Back to feed

b8578

Mar 29, 2026
Meta/llama.cppCLIvb8578

hexagon: dma optimizations (mostly fixing regressions) (#21137)

  • hex-fa: add simple dma cache for Mask

I noticed that we were refetch the mask rows over and over. This simple cache avoids that.

  • hex-dma: unset in-order desc bit which caused signficant perf regression

We don't rely on true in order processing of the DMA descriptors anywhere. Turns out this mode caused significant regression of around 3-4 TPS during token gen.

  • hex-rope: update comment to clarify that we don't need in-order DMA completions

macOS/iOS:

Linux:

Windows:

openEuler: