Why build another engine?
llama.cpp is excellent and remains a supported option. Its generality also means it carries cost for many paths that a single-user, latency-sensitive app does not need.
The custom engine will become tinyinference's default once it matches or exceeds llama.cpp on the workloads that matter to users. llama.cpp remains available for users who prefer it.
Design priorities
- Hand-written kernels
- Architecture-specific matmul and attention kernels keep the critical path close to the hardware.
- Zero-copy memory layout
- Weights, KV cache, and intermediates share a cache-aware arena to reduce unnecessary movement.
- Mmap-first by default
- Model weights stay on disk and are mapped on demand so oversized GGUF models can still run.
- Quantization-native
- Integer and blockwise quantization are handled directly in the kernel path.
- Persistent KV cache
- Session state stays resident across requests for smoother multi-turn chat.
- Small native binary
- The engine runs directly within the app with no JVM, Python process, or container.
How it gets built
- Measure first
- Every optimization starts with a profiling target and ends with a regression check.
- Use the lowest layer that wins
- SIMD, assembly, and unsafe layout are used only where measurement shows a clear benefit.
- Optimize for latency
- The goal is faster first tokens and smoother streaming on the machine in front of you.
- Ship one binary
- CPU features are detected at startup so users do not need separate builds or tuning flags.
Where it fits in the roadmap
- Today
- tinyinference uses llama-server from llama.cpp as the current default while the custom engine is developed.
- Next
- A from-scratch engine is being written alongside it, sharing the GGUF and mmap foundation.
- Later
- The custom engine becomes the default once it matches or beats llama.cpp on the workloads that matter to users. llama.cpp remains available as a supported option.