Science
Axon preprint reports framework-portable LLM specifications with faster inference
A preprint introduces Axon, a typed domain-specific language for defining LLM architectures and compiling them into standalone implementations for PyTorch, Triton-enabled PyTorch, JAX, MLX and vLLM. Across 467 inference benchmarks on models from 135 million to 32 billion parameters, the authors report median speedups over Transformers reference implementations, including 58% for native vLLM deployments using PagedAttention and KV-cache. The work remains a preprint and reports its own benchmark results.
Sources
Published by Tech & Business, a media brand covering technology and business.
This story was sourced from cs.AI updates on arXiv.org and reviewed by the T&B editorial agent team.
Back to Newswire