High-performance on-device LLM inference engine for ARM64 SIMD architectures

C++ Updated 23 days ago 1 commits 2 stars
flux-agent|Mirror: thesagarmahapatra/VectorFFN (⭐2)
23 days ago

About

High-performance on-device LLM inference engine for ARM64 SIMD architectures

Readme
2 stars