The general-purpose heap allocator — malloc/free — is a marvel of engineering.
It handles allocations of arbitrary size and lifetime, threads safely, and minimizes fragmentation across
wildly unpredictable usage patterns. It is also, for a large class of programs, completely overkill.
If you know something about your allocation pattern, you can do much better. Arena allocators are
the simplest and most powerful expression of that idea.
The core idea
An arena (also called a linear allocator or bump allocator) is a block of memory with a single cursor. Allocating means moving the cursor forward. Freeing means... nothing. You free the entire arena at once when you're done with everything in it.
typedef struct {
uint8_t *base;
size_t offset;
size_t capacity;
} Arena;
Arena arena_create(size_t capacity) {
return (Arena){
.base = malloc(capacity),
.offset = 0,
.capacity = capacity,
};
}
void *arena_alloc(Arena *a, size_t size) {
// Align to 8 bytes
size = (size + 7) & ~7;
assert(a->offset + size <= a->capacity);
void *ptr = a->base + a->offset;
a->offset += size;
return ptr;
}
void arena_reset(Arena *a) {
a->offset = 0;
}
void arena_destroy(Arena *a) {
free(a->base);
}
That's the whole allocator. Thirty lines. arena_alloc is a pointer addition and a size
increment — it cannot be faster. And because you reset the whole arena at once, there's no tracking
of individual allocations, no free list, no fragmentation, no use-after-free bugs.
Why this is so fast
A malloc call needs to find a free block of the right size, update its internal bookkeeping,
potentially call into the OS for more memory, and handle thread synchronization. A bump allocation
is pointer arithmetic. The difference in hot paths is enormous — often 10–50x faster in benchmarks,
and that's before you account for cache effects.
Cache is the real win. When you allocate a bunch of objects from an arena in sequence, they all live
next to each other in memory. When you iterate over them, you're walking contiguous bytes. The prefetcher
loves this. Compare to a list of separately malloc'd objects scattered across the heap,
and you can see why arena-allocated data structures can be dramatically faster even independent of
the allocation cost itself.
The lifetime model
Arenas make lifetime management explicit and simple. Instead of tracking which individual objects are alive, you track which phase of work is alive. This maps cleanly onto how most programs actually behave:
- Parse a file → allocate into arena → process the result → reset the arena.
- Handle a web request → arena for the duration of the request → reset on response.
- Run a game frame → frame arena for scratch allocations → reset at the end of the frame.
In each case, everything allocated together can be freed together. This is the pattern. If your problem fits it, an arena will simplify your code and speed it up simultaneously.
Scratch arenas
A particularly useful pattern is the scratch arena — a temporary arena for short-lived intermediate allocations inside a function. You allocate a chunk, do some work, and reset when done. Because reset is O(1), you can do this thousands of times per second without concern.
// Permanent arena lives for the whole program.
// Scratch arena is reset frequently for temp work.
char *build_path(Arena *scratch, const char *dir, const char *file) {
size_t saved = scratch->offset;
size_t len = strlen(dir) + strlen(file) + 2;
char *buf = arena_alloc(scratch, len);
snprintf(buf, len, "%s/%s", dir, file);
// Caller resets the scratch arena when done
return buf;
}
You can make this even cleaner by saving and restoring the arena offset in a scoped fashion, giving you something close to stack semantics with heap memory.
When arenas don't fit
Arenas aren't for everything. They break down when you need objects with individually varying lifetimes — when some things in the arena need to outlive others. They also don't help if you have a graph or tree structure where you need to delete nodes mid-traversal.
For those cases you can combine strategies: use an arena for the bulk of allocations with the same lifetime, and use a pool allocator or the general heap for things with irregular lifetimes. Most programs have a majority of allocations that fall into predictable lifetime groups — those are the ones to arena.
In practice
Once you start using arenas you start seeing them everywhere. The compiler toolchain world has used them forever — clang, LLVM, GCC all have their own arena-style allocators because parsing and IR construction create enormous numbers of short-lived nodes with exactly this structure. Game engines have used per-frame arenas as a matter of course for decades.
The general-purpose allocator is the right tool when you don't know your allocation pattern. When you do — and surprisingly often, you do — an arena will outperform it in every metric that matters: speed, memory overhead, and code simplicity.