Inference engine

A custom inference engine for every kind of hardware.

A custom LLM inference engine, built for low-level efficiency and all types of hardware.

Why build another engine?

llama.cpp is excellent and remains a supported option. Its generality also means it carries cost for many paths that a single-user, latency-sensitive app does not need.

The custom engine will become tinyinference's default once it matches or exceeds llama.cpp on the workloads that matter to users. llama.cpp remains available for users who prefer it.

Design priorities

Hand-written kernels
Architecture-specific matmul and attention kernels keep the critical path close to the hardware.
Zero-copy memory layout
Weights, KV cache, and intermediates share a cache-aware arena to reduce unnecessary movement.
Mmap-first by default
Model weights stay on disk and are mapped on demand so oversized GGUF models can still run.
Quantization-native
Integer and blockwise quantization are handled directly in the kernel path.
Persistent KV cache
Session state stays resident across requests for smoother multi-turn chat.
Small native binary
The engine runs directly within the app with no JVM, Python process, or container.

How it gets built

Measure first
Every optimization starts with a profiling target and ends with a regression check.
Use the lowest layer that wins
SIMD, assembly, and unsafe layout are used only where measurement shows a clear benefit.
Optimize for latency
The goal is faster first tokens and smoother streaming on the machine in front of you.
Ship one binary
CPU features are detected at startup so users do not need separate builds or tuning flags.

Where it fits in the roadmap

Today
tinyinference uses llama-server from llama.cpp as the current default while the custom engine is developed.
Next
A from-scratch engine is being written alongside it, sharing the GGUF and mmap foundation.
Later
The custom engine becomes the default once it matches or beats llama.cpp on the workloads that matter to users. llama.cpp remains available as a supported option.