← back to writing

Jan 2026

Arena allocators

~10 min read

The general-purpose heap allocator — malloc/free — is a marvel of engineering. It handles allocations of arbitrary size and lifetime, threads safely, and minimizes fragmentation across wildly unpredictable usage patterns. It is also, for a large class of programs, completely overkill. If you know something about your allocation pattern, you can do much better. Arena allocators are the simplest and most powerful expression of that idea.

The core idea

An arena (also called a linear allocator or bump allocator) is a block of memory with a single cursor. Allocating means moving the cursor forward. Freeing means... nothing. You free the entire arena at once when you're done with everything in it.

typedef struct {
    uint8_t *base;
    size_t   offset;
    size_t   capacity;
} Arena;

Arena arena_create(size_t capacity) {
    return (Arena){
        .base     = malloc(capacity),
        .offset   = 0,
        .capacity = capacity,
    };
}

void *arena_alloc(Arena *a, size_t size) {
    // Align to 8 bytes
    size = (size + 7) & ~7;
    assert(a->offset + size <= a->capacity);
    void *ptr = a->base + a->offset;
    a->offset += size;
    return ptr;
}

void arena_reset(Arena *a) {
    a->offset = 0;
}

void arena_destroy(Arena *a) {
    free(a->base);
}

That's the whole allocator. Thirty lines. arena_alloc is a pointer addition and a size increment — it cannot be faster. And because you reset the whole arena at once, there's no tracking of individual allocations, no free list, no fragmentation, no use-after-free bugs.

Why this is so fast

A malloc call needs to find a free block of the right size, update its internal bookkeeping, potentially call into the OS for more memory, and handle thread synchronization. A bump allocation is pointer arithmetic. The difference in hot paths is enormous — often 10–50x faster in benchmarks, and that's before you account for cache effects.

Cache is the real win. When you allocate a bunch of objects from an arena in sequence, they all live next to each other in memory. When you iterate over them, you're walking contiguous bytes. The prefetcher loves this. Compare to a list of separately malloc'd objects scattered across the heap, and you can see why arena-allocated data structures can be dramatically faster even independent of the allocation cost itself.

The lifetime model

Arenas make lifetime management explicit and simple. Instead of tracking which individual objects are alive, you track which phase of work is alive. This maps cleanly onto how most programs actually behave:

In each case, everything allocated together can be freed together. This is the pattern. If your problem fits it, an arena will simplify your code and speed it up simultaneously.

Scratch arenas

A particularly useful pattern is the scratch arena — a temporary arena for short-lived intermediate allocations inside a function. You allocate a chunk, do some work, and reset when done. Because reset is O(1), you can do this thousands of times per second without concern.

// Permanent arena lives for the whole program.
// Scratch arena is reset frequently for temp work.

char *build_path(Arena *scratch, const char *dir, const char *file) {
    size_t saved = scratch->offset;

    size_t len = strlen(dir) + strlen(file) + 2;
    char *buf = arena_alloc(scratch, len);
    snprintf(buf, len, "%s/%s", dir, file);

    // Caller resets the scratch arena when done
    return buf;
}

You can make this even cleaner by saving and restoring the arena offset in a scoped fashion, giving you something close to stack semantics with heap memory.

When arenas don't fit

Arenas aren't for everything. They break down when you need objects with individually varying lifetimes — when some things in the arena need to outlive others. They also don't help if you have a graph or tree structure where you need to delete nodes mid-traversal.

For those cases you can combine strategies: use an arena for the bulk of allocations with the same lifetime, and use a pool allocator or the general heap for things with irregular lifetimes. Most programs have a majority of allocations that fall into predictable lifetime groups — those are the ones to arena.

In practice

Once you start using arenas you start seeing them everywhere. The compiler toolchain world has used them forever — clang, LLVM, GCC all have their own arena-style allocators because parsing and IR construction create enormous numbers of short-lived nodes with exactly this structure. Game engines have used per-frame arenas as a matter of course for decades.

The general-purpose allocator is the right tool when you don't know your allocation pattern. When you do — and surprisingly often, you do — an arena will outperform it in every metric that matters: speed, memory overhead, and code simplicity.